Skip to content
← Back

garbled-text-diagnoser

Text
Runs locally · Not uploaded

This tool runs entirely in your browser. Open DevTools → Network panel and search your input — it appears in no request.

Usage Guide

A CSV export full of "ä½ å¥½", logs bristling with "©", API payloads studded with %E4%B8%AD — mojibake is bytes decoded with the wrong encoding, and almost every tool out there is a converter: it saves you only if you guess the source encoding right. This tool diagnoses: paste the garbled text, it reads the byte signatures, tells you what happened and how to fix it.

1. Six recognizable failure modes

UTF-8 read as GBK (the ä½ å¥½ shape: runs of Latin-extended pairs); UTF-8 read as Latin-1 ( + symbol, like ©); double-encoded UTF-8 (é — someone fixed it in the wrong direction); replacement-char runs (����, already lossy, the one unrecoverable mode); percent-encoding residue (%E4%B8%AD, one URL-decode away from health); \u or HTML-entity residue (中 / 中, serialization never unwrapped). Each mode has a different fix — diagnose first, then act; it beats trying encodings one by one.

2. Why mojibake happens (one page)

Text is bytes in storage and transit. "你好" is six bytes in UTF-8; read as GBK, every two bytes become one character — yielding three strange glyphs. Read GBK bytes as UTF-8 and the multi-byte structure collapses into replacement characters. The core law: **UTF-8 misread by a single-byte encoding (GBK/Latin-1) is recoverable; single-byte misread as UTF-8 that produced U+FFFD is gone for good.** That is exactly the line this tool draws between "fixable" and "lost".

Related Tools

Related Articles