HTML Entity Encoder
Encode text for HTML or decode entities back to characters — the five markup breakers, the named Latin-1 and symbol sets, and numeric references up to emoji. Decoding runs exactly once, so text about HTML survives untouched.
The entities worth recognising
| Entity | Character | Where it appears |
|---|---|---|
& < > | & < > | the three every template engine must write |
" ' | " ' | inside attribute values |
| non-breaking space | keeping "10 kg" on one line |
— – | — – | typographic dashes |
© ™ € | © ™ € | legal and currency marks |
🙂 | 🙂 | anything without a name, by code point |
Everything runs in your browser; nothing you paste is sent anywhere.
Frequently asked questions
Which characters does minimal encoding touch?
Exactly five: & becomes &, < and > become < and >, and the two quote characters become " and '. Those are the characters that can break out of markup or an attribute — everything else, including accented letters and emoji, is perfectly legal in an HTML document as-is when the page declares UTF-8.
When would I need the full mode?
When the text will pass through a system that mangles bytes outside ASCII — a legacy email pipeline, an old CMS, a config file with unclear encoding. Full mode writes every non-ASCII character as a named entity where one exists (é, —) and a hex reference (ə) where none does, producing pure-ASCII output that survives any transport.
Why does the decoder leave &lt; as <?
Because decoding must run exactly once. &lt; is the encoding of the four characters < — someone writing documentation about HTML. A decoder that keeps going until nothing changes would turn that into <, corrupting every document that talks about markup. If your text really was encoded twice, run decode twice, deliberately.
Why did AT&T; pass through unchanged?
There is no entity named T, so &T; is not a reference — it is two characters of prose that happen to look like one. Eating it would invent text that was never there. The same applies to bare ampersands: a & b needs no decoding, and the tool does not guess.
What are the A and 🙂 forms?
Numeric character references: decimal after &#, hexadecimal after &#x, naming a Unicode code point directly. They cover all of Unicode — 🙂 is 🙂 — which is why full encoding falls back to them for characters without a name. References to impossible code points (beyond U+10FFFF, or lone surrogates) are left untouched rather than replaced with garbage.
Is this the same as URL encoding?
No — different layer, different syntax. URL percent-encoding (%20, %C3%A9) protects characters inside web addresses; HTML entities protect characters inside markup. A URL inside an HTML attribute may legitimately need both, applied in the right order: percent-encode the URL first, then entity-encode the attribute. The URL encoder on this site handles the other half.
Related tools
- Diagram & Flowchart MakerDraw flowcharts and diagrams, connect the boxes with arrows that route themselves, and let it arrange the whole thing. Exports SVG. Nothing is uploaded.
- Password GeneratorGenerate strong random passwords using your browser's cryptographic RNG, with a live strength estimate.
- Regex TesterTest regular expressions with live highlighting, groups and replace preview — JavaScript's own engine, with a kill switch for catastrophic patterns.