HTML Entity Encoder

Works offlineNothing is uploadedFree, no sign-up

Encode text for HTML or decode entities back to characters — the five markup breakers, the named Latin-1 and symbol sets, and numeric references up to emoji. Decoding runs exactly once, so text about HTML survives untouched.

The entities worth recognising

EntityCharacterWhere it appears
&amp; &lt; &gt;& < >the three every template engine must write
&quot; &apos;" 'inside attribute values
&nbsp;non-breaking spacekeeping "10 kg" on one line
&mdash; &ndash;— –typographic dashes
&copy; &trade; &euro;© ™ €legal and currency marks
&#x1F642;🙂anything without a name, by code point

Everything runs in your browser; nothing you paste is sent anywhere.

Frequently asked questions

Which characters does minimal encoding touch?

Exactly five: & becomes &amp;, < and > become &lt; and &gt;, and the two quote characters become &quot; and &apos;. Those are the characters that can break out of markup or an attribute — everything else, including accented letters and emoji, is perfectly legal in an HTML document as-is when the page declares UTF-8.

When would I need the full mode?

When the text will pass through a system that mangles bytes outside ASCII — a legacy email pipeline, an old CMS, a config file with unclear encoding. Full mode writes every non-ASCII character as a named entity where one exists (&eacute;, &mdash;) and a hex reference (&#x259;) where none does, producing pure-ASCII output that survives any transport.

Why does the decoder leave &amp;lt; as &lt;?

Because decoding must run exactly once. &amp;lt; is the encoding of the four characters &lt; — someone writing documentation about HTML. A decoder that keeps going until nothing changes would turn that into <, corrupting every document that talks about markup. If your text really was encoded twice, run decode twice, deliberately.

Why did AT&T; pass through unchanged?

There is no entity named T, so &T; is not a reference — it is two characters of prose that happen to look like one. Eating it would invent text that was never there. The same applies to bare ampersands: a & b needs no decoding, and the tool does not guess.

What are the &#65; and &#x1F642; forms?

Numeric character references: decimal after &#, hexadecimal after &#x, naming a Unicode code point directly. They cover all of Unicode — &#x1F642; is 🙂 — which is why full encoding falls back to them for characters without a name. References to impossible code points (beyond U+10FFFF, or lone surrogates) are left untouched rather than replaced with garbage.

Is this the same as URL encoding?

No — different layer, different syntax. URL percent-encoding (%20, %C3%A9) protects characters inside web addresses; HTML entities protect characters inside markup. A URL inside an HTML attribute may legitimately need both, applied in the right order: percent-encode the URL first, then entity-encode the attribute. The URL encoder on this site handles the other half.

Related tools