Unicode Inspector
See what a string is actually made of: every code point named — CYRILLIC SMALL LETTER A, ZERO WIDTH SPACE, EGYPTIAN HIEROGLYPH A001 — with the block it comes from, its UTF-8 and UTF-16 bytes, Cyrillic lookalikes flagged, and the four different lengths — code points, UTF-16 units, UTF-8 bytes, graphemes — that explain why the same text 'has three lengths' in three systems.
Four rulers for one string
Unicode separates what a character is (a code point) from how it is stored (UTF-8 or UTF-16 bytes) and from what a reader sees (a grapheme). Most string bugs are two of those rulers being confused: a database column limited in bytes rejecting a text counted in characters, a substring cut in the middle of a surrogate pair, a cursor stepping through half a flag.
The table view is deliberately byte-level: seeing C9 99 next to ə teaches more about UTF-8 than a chapter of explanation — and seeing 200B next to nothing at all usually ends a two-hour debugging session.
The name and the note are different questions
The Name column answers what a character is; the Note column answers why you should care about this one. They are kept apart on purpose. A name is not a warning — LATIN SMALL LETTER A and CYRILLIC SMALL LETTER A are equally respectable names, and only the second is a reason to look twice at a domain. So the note carries the part the name cannot: that an en dash is not a hyphen, that U+FEFF is also a byte order mark, that this letter is a lookalike. Where the note would only repeat the name, it steps aside and marks the row as a suspect instead.
Frequently asked questions
Why does my string have three different lengths?
Because systems count different things: JavaScript's .length counts UTF-16 units (an emoji is 2), databases usually meter UTF-8 bytes (ə is 2, a flag emoji 8), and what a user calls one character is a grapheme, which may be many code points. This page prints all four so the mismatch stops being a mystery.
What is a zero-width space and how did it get in my text?
U+200B is a legal line-break hint with no ink — copy a snippet from a formatted web page or a chat app and it rides along. It then breaks equality checks, logins and SQL matches while looking like nothing. The suspects counter exists mostly for this one character.
What is a Cyrillic lookalike attack?
Letters like а, е, о, с render pixel-identical to Latin a, e, o, c but are different code points. A domain or username spelled with them passes visual inspection and fails every string comparison — which is exactly how homograph phishing works. The inspector names each such letter.
Why is é sometimes one code point and sometimes two?
Unicode allows both: precomposed U+00E9, or e plus a combining accent. They render identically and compare unequal — the classic reason a filename search misses an existing file. Normalisation (NFC) fixes comparisons; this page shows you which form you are holding.
Where do the character names come from?
From the Unicode Character Database — the same file the standard publishes, carried with the page rather than fetched from an API, so the names work offline and nothing about your text is sent anywhere. 40,470 code points are listed by name; a further 119,000 — Chinese ideographs, Korean syllables, Tangut — have names the standard builds by rule, and so does this page. Private use and surrogate code points have no name at all: the standard gives them none, so their Name cell stays empty rather than being filled with a guess. The version and its publication date are printed under the table, because an empty Name can also mean the character was unassigned when that snapshot was taken.
Why do the names appear a moment after the table?
The name table is half a megabyte. It is loaded as a separate file the first time you open this tool, so every other page on the site stays light and the table you are reading paints immediately with the columns that need no lookup. After the first visit your browser has it cached, and the tool keeps naming characters with no network at all.
Related tools
- Word & Character CounterLive word, character, sentence and paragraph counts, with reading time, keyword density, and whether the text fits a post, a search title or one text message.
- Sort & Dedupe LinesSort, dedupe, shuffle, trim and number lines — natural sort puts file2 before file10 — or compare two lists.
- Lorem Ipsum GeneratorPlaceholder text by paragraphs, sentences or exact word count — starting with the classic line everyone recognizes, shaped like real prose.