Access succeeded. Extraction did not.
A crawler receives a page. The response is successful, the paragraph exists, and the extracted sentence is grammatical. Yet it says something different from the paragraph a person sees.
This is a different failure from being blocked. The request succeeded, but the representation being collected is wrong. Font-based text obfuscation creates a gap between encoded characters and rendered glyphs. An extraction pipeline can preserve every source character and still fail to preserve what the reader understood.
This article examines that gap. It is a technical analysis, not a benchmark of provider effectiveness or a claim about AI training outcomes.
The missing layer: text is not a picture of text
HTML contains characters. Fonts supply glyphs, and text shaping can replace glyphs before they reach the screen. The OpenType character map associates character codes with glyph indices. Glyph substitution can substitute one glyph or a sequence of glyphs during layout.
Those mechanisms serve ordinary typography and writing systems. Their presence does not indicate obfuscation. But they can also be used to make an encoded sequence look like a different word. The browser’s rendered output changes while the underlying character sequence remains unchanged.
That is why merely switching from an HTTP client to a browser is not necessarily enough. A browser may render the correct-looking page while its DOM text APIs still return the encoded sequence. Reading text nodes is a different measurement from reading rendered pixels.
A real example: ShieldFont
ShieldFont is an open-source project that uses word substitutions and custom fonts to make source text differ from the writing displayed on screen. Its stated aim is to increase the cost of unauthorized bulk collection for AI training, rather than make targeted recovery permanently impossible.
This is one implementation of the broader extraction challenge, not a synonym for all font obfuscation. Its current public materials also describe custom mappings and an alternative reader-access mechanism. The implementation is evolving, so a result for one font or fixture should not be presented as coverage of every deployment. See the project repository for its current scope and qualifications.
Methodology
This analysis combines a review of OpenType’s character/glyph distinction, ShieldFont’s public description, and the documented scope of our ShieldFont Decoder project. It does not repeat the provider’s benchmark or measure its current live site.
The question is narrower: which representation does an extractor collect, and what evidence would establish that the intended text was recovered? The comparison below describes representations rather than observed accuracy rates.
| Representation | What it records | What it does not establish |
|---|---|---|
| HTML or DOM text | Stored character sequences | That those characters match the words drawn on screen |
| Screenshot | Rendered pixels in one environment | Exact Unicode, correct reading order, or coverage of unseen content |
| Font metadata | Available mappings and shaping rules | That every visible substitution is reversibly described |
| Decoded output | A recovery method’s interpretation | Correctness without an expected result and coverage checks |
Findings
A plausible sentence can be a bad result
Extraction validation often checks whether text is present, readable, and long enough. These checks can accept fluent substitutions. The failure is semantic: a name, quantity, relationship, or claim can change while the sentence remains structurally ordinary.
A valid-result definition therefore needs more than a nonempty string. In a controlled fixture, compare the output with the intended text. On unfamiliar content, treat disagreement between representations as an investigation signal rather than assuming whichever string looks fluent is correct.
Rendering and character extraction are separate paths
A successful browser load confirms that a page can be rendered in that environment. It does not confirm that textContent, innerText, or copied text recovers the glyphs’ apparent meaning. Site scripts and explicit plain-text modes can change the DOM or clipboard behavior, so the tested extraction path must be recorded.
The font can be evidence
Our decoder takes a font-aware approach for supported reversible schemes. It reconstructs a mapping from OpenType substitution, composite-glyph, and character-map information and applies it to the encoded text. It does not use OCR.
The implementation and its local workflow are documented in ShieldFont Decoder: Recovering Font-Obfuscated Text. That page is the home for the tool. This article is the home for the underlying extraction problem.
Recovery depends on the actual font and scheme. A mapping that works for a bundled fixture is not evidence that another font, a new mapping, or the latest provider release is supported.
Limitations
- This is a source-based analysis and an extraction model, not a measured accuracy comparison.
- Ordinary ligatures and script shaping are legitimate typography. Custom fonts alone are not evidence of hostile content.
- Recoverable mappings are scheme-specific. Glyph outlines, contextual rules, or transformations without suitable metadata may require another approach.
- Pixel-based recovery has separate reading-order, layout, and recognition concerns. No OCR accuracy or cost claim is made here.
- Accessibility, clipboard behavior, and search indexing can vary by integration. Our decoder does not validate those features or the provider’s AI-training claims.
- ScrapeTrace and its decoder are independent of the ShieldFont project.
Reproduction
Use an owned fixture with a font you are permitted to analyze. Preserve four things: the intended paragraph, encoded source characters, the exact font file, and the browser-rendered result. Record the browser and extraction API used.
Compare source extraction with the intended paragraph before applying any recovery method. For a supported font, follow the decoder’s local setup, then compare its output character-for-character with your expected text. Report unchanged words, unresolved tokens, punctuation, and partial coverage alongside successes.
No new experiment is claimed by this suggested protocol. Its purpose is to make a future result auditable: request success, readable rendering, and correct extraction are three separate outcomes.
Found a mistake or a result you cannot reproduce?
Send a correctionReference: /articles/font-based-text-obfuscation/