Projects / Tool

ShieldFont Decoder: Recovering Font-Obfuscated Text

Explore why source text can differ from rendered content, and how ShieldFont Decoder recovers copyable Unicode from reversible OpenType font mappings.

The extraction problem

A successful page request does not guarantee correct text. With font-based obfuscation, characters stored in the DOM can differ from the words that a browser renders. Reading the source alone can produce plausible-looking but incorrect extraction.

ShieldFont Decoder investigates this gap by analyzing the font metadata responsible for the transformation. Its output is normal, copyable Unicode text. It uses OpenType mappings rather than OCR or browser automation.

General approach

The decoder reads substitution and composite-glyph information from the font, relates that information to its Unicode character map, and builds a reusable source-to-rendered mapping. It applies that mapping to the obfuscated text using word boundaries appropriate to the supported scheme.

The implementation uses Python, Flask, and fontTools. It supports WOFF2, WOFF, TTF, and OTF font uploads. Extracted mappings are cached in memory for the server lifetime, and a pre-generated mapping for the bundled default font is included.

Try the project

The ShieldFont Decoder demo starts with the bundled font mapping. Its Advanced section can analyze another font. The GitHub repository contains the implementation and setup instructions.

For a controlled comparison, use your own fixture: retain its source text, font, expected rendered text, and decoded output. Those artifacts let you check whether the recovered mapping matches the intended content without inferring accuracy from a successful request.

What this demonstrates

Extraction reliability can fail after access succeeds. In this project, the font is part of the content transformation, so analyzing only DOM text misses the relevant layer.

The project documentation describes recovery from reversible substitution schemes. This page is a tool overview, not a new accuracy benchmark or a claim that every obfuscated font can be decoded.

Limitations

  • The supported approach needs recoverable OpenType substitution and composite-glyph metadata, notably the GSUB and glyf information used by the target scheme.
  • Different glyph transformations or missing mappings may require another analysis method.
  • Supported file formats do not imply support for every obfuscation scheme inside those files.
  • Font metadata recovery does not establish completeness or correctness for an entire scraped page.
  • The demo is a separate Flask application. ScrapeTrace itself remains a static publication.

Reproduction

Run the project locally with a font and text fixture you are permitted to inspect:

git clone https://github.com/Logesh08/shieldfont-decoder.git
cd shieldfont-decoder
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python app.py

The local application listens on port 5000. On Windows, activate the virtual environment with .venv\Scripts\activate. Follow the project README for current requirements.

Independence and source

ShieldFont Decoder is an independent research and compatibility tool. It is not affiliated with or endorsed by the ShieldFont project.

ShieldFont Decoder on GitHub · Created by Logesh Krishna.

Found a mistake or a result you cannot reproduce?

Send a correctionReference: /projects/shieldfont-decoder/

Research capture

100%