Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

HTML Entity Encoder & Decoder

Escape text for HTML, or decode &, © and → exactly as a browser would.

Developer No upload Works offline Free, no sign-up

HTML entity list

All 2,231 named references of the HTML Standard. Search by name, character, code point or meaning.

58 common entities

  •   U+00A0 ·   no-break space
  • & U+0026 · & ampersand
  • < U+003C · < less-than sign
  • > U+003E · > greater-than sign
  • " U+0022 · " quotation mark
  • ' U+0027 · ' apostrophe
  • © U+00A9 · © copyright sign
  • ® U+00AE · ® registered sign
  • ™ U+2122 · ™ trade mark sign
  • € U+20AC · € euro sign
  • £ U+00A3 · £ pound sign
  • ¥ U+00A5 · ¥ yen sign
  • ¢ U+00A2 · ¢ cent sign
  • ° U+00B0 · ° degree sign
  • ± U+00B1 · ± plus-minus sign
  • × U+00D7 · × multiplication sign
  • ÷ U+00F7 · ÷ division sign
  • − U+2212 · − minus sign
  • ≠ U+2260 · ≠ not equal to
  • ≤ U+2264 · ≤ less-than or equal to
  • ≥ U+2265 · ≥ greater-than or equal to
  • ∞ U+221E · ∞ infinity
  • µ U+00B5 · µ micro sign
  • ¶ U+00B6 · ¶ pilcrow sign
  • § U+00A7 · § section sign
  • · U+00B7 · · middle dot
  • • U+2022 · • bullet
  • … U+2026 · … horizontal ellipsis
  • – U+2013 · – en dash
  • — U+2014 · — em dash
  • ‘ U+2018 · ‘ left single quotation mark
  • ’ U+2019 · ’ right single quotation mark
  • “ U+201C · “ left double quotation mark
  • ” U+201D · ” right double quotation mark
  • « U+00AB · « left-pointing double angle quotation mark
  • » U+00BB · » right-pointing double angle quotation mark
  • ← U+2190 · ← leftwards arrow
  • → U+2192 · → rightwards arrow
  • ↑ U+2191 · ↑ upwards arrow
  • ↓ U+2193 · ↓ downwards arrow
  • ↔ U+2194 · ↔ left right arrow
  • ⇒ U+21D2 · ⇒ rightwards double arrow
  • ✓ U+2713 · ✓ check mark
  • ✗ U+2717 · ✗ ballot x
  • ☆ U+2606 · ☆ white star
  • ★ U+2605 · ★ black star
  • ♥ U+2665 · ♥ black heart suit
  • ½ U+00BD · ½ vulgar fraction one half
  • ¼ U+00BC · ¼ vulgar fraction one quarter
  • ¾ U+00BE · ¾ vulgar fraction three quarters
  • ² U+00B2 · ² superscript two
  • ³ U+00B3 · ³ superscript three
  • ­ U+00AD · ­ soft hyphen
  • ‍ U+200D · ‍ zero width joiner
  • ‌ U+200C · ‌ zero width non-joiner
  •   U+2009 ·   thin space
  •   U+2003 ·   em space
  •   U+2002 ·   en space

Source: WHATWG HTML Standard — named character references; character names from the Unicode Character Database.

Next steps

About the HTML Entity Encoder & Decoder

Paste text to escape it for HTML — only the five special characters (&, <, >, ", '), or those plus every non-ASCII character — as names (&copy;), decimal (&#169;) or hex (&#xA9;) references. Emoji and other characters without a name get numeric references automatically.

Paste HTML to decode it: all 2,231 named character references of the WHATWG HTML Standard plus numeric ones, with the rules browsers use — the longest name wins, legacy names such as &copy work without a semicolon, attribute values keep &copy= as typed, and out-of-range numbers become U+FFFD. Every reference that a browser would flag is listed with the reason. A searchable list of every entity, with copy buttons, is below the converter.

How to use it

  1. Choose Encode or Decode and paste your text or HTML. The result updates as you type.
  2. When encoding, choose what to encode (only the special characters, everything outside ASCII, or every character) and whether to use names, decimal or hex.
  3. When decoding an href or another attribute value, tick Use attribute-value rules; for text that was escaped twice (&amp;lt;), tick Also decode double-encoded text.
  4. Copy or download the result, or press Use as input to convert it back.
  5. Search the entity list by name (rarr), character (©), code point (U+2192) or meaning (arrow, copyright) and copy the entity or the character.

Examples

Escape a snippet for display in HTML
Input
<a href="x">Tom & Jerry's</a>
Result
&lt;a href=&quot;x&quot;&gt;Tom &amp; Jerry&apos;s&lt;/a&gt;
Names where they exist, numbers for the rest
Input
café → 😀
Result
caf&eacute; &rarr; &#128512;

With Hex instead: caf&#xE9; &#x2192; &#x1F600;.

Decode the way a browser does
Input
&notit; &notin; &copy MySmartCoPilot
Result
¬it; ∉ © MySmartCoPilot

&notit; is not a name, but &not is a legacy name that works without a semicolon — the HTML Standard’s own example. The missing semicolons are reported.

A URL in an attribute
Input
/search?a=1&copy=2
Result
Attribute rules: /search?a=1&copy=2
Text rules:      /search?a=1©=2
Double-encoded text
Input
&amp;lt;b&amp;gt;
Result
one pass: &lt;b&gt; · fully decoded: <b>

Common uses

  • Showing code samples or user-supplied text on a web page without the browser treating it as markup.
  • Cleaning up text exported from a CMS, an RSS feed or an API that arrives full of &amp;, &quot; and &#8217;.
  • Finding the entity for a symbol — non-breaking space, ©, ™, €, arrows, maths operators, Greek letters — and copying it.
  • Writing HTML email or other ASCII-only output, where every non-ASCII character has to be a reference.
  • Debugging double-escaping bugs and URLs whose & parameters were mangled.

Which characters need escaping?

In HTML text, escape & and < (and, by convention, >). Inside an attribute value, also escape the quote character that delimits it: &quot; for ", and &#39; or &apos; for '. Characters outside ASCII do not need escaping on a UTF-8 page (<meta charset="utf-8">), but references keep them intact in ASCII-only channels such as some email systems.

Escaping is not the same as sanitising. HTML entity encoding protects text and quoted attributes, but inside <script>, <style>, URLs and event-handler attributes other encodings apply (see the OWASP Cross Site Scripting Prevention Cheat Sheet), and to allow some markup you need an HTML sanitiser.

How browsers decode entities

  • Names are case-sensitive and the longest match wins: &notin; is ∉, while &notit; becomes ¬it; because only &not matches.
  • The 106 legacy names from HTML 4 and Latin-1 (&amp, &lt, &copy, &eacute…) also work without the semicolon; all others need it.
  • In attribute values, a legacy name without a semicolon followed by = or a letter or digit is left alone, so old URLs such as ?a=1&copy=2 keep working.
  • &#0;, surrogates (&#xD800;) and values above &#x10FFFF; become U+FFFD. Numbers 128–159 are read as Windows-1252: &#128; is €, &#150; is –.

The list of 2,231 names is fixed: the HTML Standard says it “is static and will not be expanded or changed in the future”.

Named, decimal or hex?

Named references are the most readable, but only 1,446 characters have one, and emoji have none. Numeric references work for every character in every browser: decimal (&#8364;) and hex (&#x20AC;) are equivalent. When a character has several names (→ is &rarr;, &rightarrow;, &RightArrow;, &srarr; and &ShortRightArrow;), the encoder uses the classic HTML 4 name or the shortest one. &apos; is HTML5 and XML; in old HTML 4 documents use &#39;.

Limitations

  • The encoder escapes characters; it does not sanitise HTML or make untrusted input safe in scripts, styles or URLs.
  • Decoding follows the HTML rules for text and attribute values. XML only predefines &lt;, &gt;, &amp;, &quot; and &apos;, so other names in an XML file need a DTD.
  • It decodes references; it does not parse or remove tags. For that, use the HTML Formatter or an HTML sanitiser.
  • In Every character mode, line breaks are kept as they are so the layout survives.

Privacy

Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.

Frequently asked questions

What is the HTML entity for a space that does not break?

&nbsp; (U+00A0, no-break space), or &#160;. For a narrow no-break space use &#8239;; for a space that may break but has a fixed width, &ensp;, &emsp; or &thinsp;.

Do I have to escape quotes?

Only inside attribute values delimited by that quote: &quot; in title="…", and &#39; (or &apos;) in title='…'. In ordinary text, quotes are fine as they are.

Why does &copy work without a semicolon but &rarr does not?

&copy is one of 106 legacy names from HTML 4 and Latin-1 that browsers accept without the semicolon (it is still reported as an error). Newer names such as &rarr; must end with ;, otherwise the text stays as typed.

Why does &#150; show a dash instead of a control character?

For historical reasons, browsers read numeric references 128–159 as Windows-1252 characters: &#150; is –, &#128; is €, &#153; is ™. The decoder does the same and tells you the correct code point (&#8211; for –).

How do I decode text that was escaped twice?

Text such as &amp;lt;b&amp;gt; was escaped twice, so one pass gives &lt;b&gt;. Tick Also decode double-encoded text (or press Decode it fully when the page notices it) to repeat until nothing is left to decode.

Is my text uploaded?

No. Encoding, decoding and searching run in your browser, and the page works offline once loaded.

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.