Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

Unicode Character Inspector

Paste text to see each character — and the invisible or look-alike ones hiding in it.

Text No upload Works offline Free, no sign-up

Everything is checked in your browser with the Unicode 18.0 data. Nothing is uploaded.

Look up a character

Names search every Unicode character name; ideographs are found by code point.

    Next steps

    About the Unicode Character Inspector

    Paste any text — a username, a domain, a line of code, a message you were asked to forward — and see every character in it: its code point (U+0041), official Unicode name, general category, script, block, bidirectional class, the Unicode version it was added in, its UTF-8 and UTF-16 bytes, and how to write it as an escape in JavaScript, Python, HTML, CSS, URLs and more.

    A security check looks for what you cannot see: zero-width and other invisible characters, bidirectional controls that can make source code display in a different order than it runs (the Trojan Source attack, CVE-2021-42574), text hidden in invisible tag characters or variation selectors, look-alike letters such as the Cyrillic а in “pаypal”, and words that mix scripts. The data comes from Unicode 18.0, and nothing you paste leaves your browser.

    How to use it

    1. Paste or type the text. The counts show characters as you see them, code points, UTF-16 units and UTF-8 bytes.
    2. Read the Security check: it lists invisible characters, bidi controls, hidden text, look-alikes and mixed-script words, with what each one is.
    3. Select a code point in the Characters table to see all its properties and escapes, each with a copy button.
    4. Use Clean copy to remove bidi controls and invisible characters (joiners that emoji and Indian or Persian words need are kept), or look up any character by its code point (U+20B9, 1F600) or by words of its name (snowman, smiling face).

    Examples

    A look-alike domain
    Input
    pаypal.com
    Result
    U+0070 LATIN SMALL LETTER P
    U+0430 CYRILLIC SMALL LETTER A (looks like a)
    U+0079 LATIN SMALL LETTER Y …

    The word mixes Latin and Cyrillic letters. Its skeleton (what it looks like) is the same as paypal.

    Two ways to write é
    Input
    café / café
    Result
    U+00E9 LATIN SMALL LETTER E WITH ACUTE
    U+0065 LATIN SMALL LETTER E + U+0301 COMBINING ACUTE ACCENT

    Both look the same, but only NFC normalization makes them equal when compared or searched.

    A Trojan Source line
    Input
    /* RLO } LRI if (isAdmin) PDI LRI begin admins only */
    Result
    Line 1: an override and an isolate are still open at the end of the line.

    Editors and code review tools then show the comment and the code in a different order than the compiler reads them.

    Common uses

    • Finding the zero-width space or no-break space that breaks a password, a URL, a CSV import or a search.
    • Checking a domain, username or sender name for look-alike letters before you trust it.
    • Reviewing source code or a pull request for bidirectional controls (Trojan Source).
    • Revealing text hidden in tag characters, as in prompt-injection tricks against AI chatbots.
    • Getting the code point, UTF-8 bytes or escape for a character you need in code, HTML or CSS.

    What the security check looks for

    • Invisible characters: everything Unicode marks as a default ignorable code point — zero-width space, word joiner, soft hyphen, byte order mark, Hangul fillers and others — plus characters such as the Braille blank that draw nothing. Zero-width joiners and non-joiners are listed separately where they do a job — in emoji sequences such as 👨‍👩‍👧 and 🏳️‍🌈, and in scripts such as Devanagari, Malayalam and Persian; between Latin, Cyrillic or Greek letters or digits, as in a “paypal” with a joiner hidden inside, they count as invisible characters.
    • Bidirectional controls (UAX #9): LRE, RLE, LRO, RLO, PDF, LRI, RLI, FSI, PDI and the marks LRM, RLM and ALM. A line where an embedding, override or isolate is still open at its end is the pattern of Trojan Source (CVE-2021-42574).
    • Hidden text: tag characters U+E0020–E007E mirror ASCII but are invisible, so a whole sentence can hide inside normal-looking text; it is decoded and shown. Runs of variation selectors after a character can hide bytes the same way. The tags of emoji subdivision flags (🏴 England, Scotland, Wales) are recognised as normal.
    • Look-alike letters and digits that can pass for Latin ones, from the Unicode confusables data (UTS #39), such as Cyrillic а for Latin a or Greek Ο for Latin O. In Russian, Greek or Hindi words they are ordinary letters, so on their own they are only noted.
    • Words that pass for Latin words: written wholly in another alphabet, every letter a look-alike, such as Cyrillic “аррӏе” for apple (UTS #39 whole-script confusables). They are flagged in domains, email addresses and user names (аррӏе.com), when they are the only word you paste, and inside Latin text.
    • Mixed-script words: words whose letters belong to no single script, using the resolved script sets of UTS #39 (Japanese mixing kanji and kana, for example, counts as one writing system).
    • Unusual spaces and controls: no-break, thin and ideographic spaces, control characters, private-use, unassigned and noncharacter code points.

    Characters, code points and bytes

    What looks like one character can be several code points: é may be one (U+00E9) or two (e + U+0301), and a family emoji is five joined by zero-width joiners. The inspector groups code points into the characters you see (grapheme clusters, UAX #29) and numbers them, so you can see which belong together. JavaScript counts UTF-16 units (an emoji counts 2), files and many databases count UTF-8 bytes (an emoji takes 4), and people count characters as they see them: all three are shown.

    The four normalization forms (UAX #15) are shown with whether your text is already in them: NFC composes characters (the usual form for storage and comparison), NFD decomposes them, and NFKC and NFKD also replace compatibility characters, so fi becomes fi and ① becomes 1.

    Limitations

    • Character names, categories and scripts follow Unicode 18.0. How a character looks depends on the fonts on your device: some may show as empty boxes.
    • Look-alike checks use the Unicode confusables list, which covers characters that look alike in common fonts; it cannot cover every font, and different words can still look similar.
    • The table lists the first 5,000 code points; the counts and the security check cover the whole text.
    • Character names (about 450 KB of data) load after the rest of the page; until then names show as “…”.

    Privacy

    Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.

    Frequently asked questions

    How do I find invisible characters in text?

    Paste the text here. The Security check lists the zero-width and other invisible characters by name, the lines with bidi controls, and how many unusual spaces there are; the Characters table marks each one where it is. Clean copy gives the text without them.

    What is the Trojan Source attack?

    Bidirectional control characters (meant for mixing Arabic or Hebrew with left-to-right text) can make code display in a different order than a compiler reads it, so a reviewer sees harmless code while something else runs. It was published as CVE-2021-42574. The inspector flags every bidi control and the lines where one is left open.

    How can I tell if a domain or username uses look-alike letters?

    Paste it here: letters from other scripts that look like Latin ones, such as Cyrillic а or Greek ο, are marked as look-alikes, words that mix scripts are listed with what they look like, and a name written wholly in another alphabet that reads as a Latin word (Cyrillic аррӏе.com for apple.com) is flagged too.

    What is the difference between a character, a code point and a byte?

    A code point is one Unicode value, written U+0041. A character, as you see it, can be several code points (an accent, an emoji with a skin tone). Bytes are how code points are stored: UTF-8 uses 1 to 4 bytes per code point, UTF-16 uses 2 or 4.

    How do I type or escape a character in code?

    Select it in the table, or look it up by code point or name, and copy the escape for your language: \u00E9 in JavaScript and JSON, \N{…} by name in Python, é in HTML or \E9 in CSS.

    Is my text uploaded?

    No. All the Unicode data is loaded into your browser and the analysis runs there; nothing you paste is sent anywhere.

    Quick answers and tool search

    Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.