Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

Text Cleaner

A checklist of clean-ups for pasted text — see what each one changed.

Text No upload Works offline Free, no sign-up

Nothing leaves your device.

Clean-ups applied in this order

  1. Keeps the visible text with sensible line breaks. Scripts, styles and comments are dropped; nothing is run or loaded.

  2. &amp; → &, &lt; → <, &nbsp; → no-break space, &#8217; → ’

  3. NFC joins letters and accents into single characters. NFKC also turns ligatures (fi), full-width letters (A) and circled numbers (①) into plain characters.

  4. Zero-width spaces, soft hyphens, byte-order marks, direction controls and other invisible characters. Joiners inside emoji and Indic or Arabic words are kept.

  5. “ ” « » → " ‘ ’ → ' – → - — → -- … → ...

  6. Including skin tones, flags and joined emoji. Plain ©, ® and ™ are kept.

  7. Links that start with http://, https://, ftp:// or www.

  8. Including a leading “mailto:”.

  9. Two or more spaces in a row become one. Runs of tabs alone are kept.

  10. Removes indentation and trailing spaces.

  11. Join hard-wrapped lines, for example text copied from a PDF or an email.

  12. Lines that are empty or contain only spaces.

Next steps

About the Text Cleaner

Text copied from web pages, PDFs, Word documents or chat apps often carries baggage: HTML tags, &amp; entities, curly quotes, emoji, links, invisible zero-width characters, double spaces and hard line wrapping. The Text Cleaner applies the clean-ups you tick, always in the same order, and shows how many changes each one made, so you know exactly what happened to your text.

HTML is converted to text by a small built-in parser: tags are removed, scripts and styles are dropped, and paragraphs, line breaks, list items and table cells become line breaks and tabs. Nothing in the HTML is run or loaded, and your text never leaves the browser.

How to use it

  1. Paste your text or HTML into the box, or use Open file (or drop a file on the box) to load one from your device.
  2. Tick the clean-ups you want. They are applied from top to bottom, in the order shown.
  3. Check the count next to each step to see what it changed, and review the result.
  4. Copy or download the cleaned text.

Examples

Plain text from an HTML snippet (Remove HTML tags + Decode HTML entities)
Input
<p>Fish &amp; chips<br>£4.50</p><p>Open <b>daily</b></p>
Result
Fish & chips
£4.50

Open daily
Curly quotes and dashes to plain ASCII
Input
“It’s ready” — she said…
Result
"It's ready" -- she said...
Unwrap text copied from a PDF (Line breaks: join lines, keep paragraphs)
Input
The quick brown
fox jumps over
the lazy dog.

A new paragraph.
Result
The quick brown fox jumps over the lazy dog.

A new paragraph.

Common uses

  • Turning an HTML email, newsletter or web page source into readable plain text.
  • Preparing copy for a CMS, spreadsheet or code that only accepts plain ASCII quotes and dashes.
  • Removing hidden zero-width and direction-control characters that break search, comparisons and code.
  • Unwrapping hard line breaks from PDFs and emails before editing or translating.
  • Stripping emoji, links or email addresses before analysing or publishing text.

Order of operations

  1. Remove HTML tags
  2. Decode HTML entities
  3. Normalise Unicode (NFC or NFKC)
  4. Remove zero-width and control characters
  5. Smart quotes and dashes → plain ASCII
  6. Remove emoji
  7. Remove URLs
  8. Remove email addresses
  9. Collapse repeated spaces
  10. Trim spaces at the start and end of lines
  11. Join lines (keeping paragraphs, or into one line)
  12. Remove blank lines

The order matters. Tags are removed before entities are decoded, so an encoded &lt;b&gt; in your text ends up as the visible text <b> instead of being treated as a tag. Space and line clean-ups come last, so they also tidy the gaps left by removed emoji, links and tags.

Which invisible characters are removed

Zero-width spaces (U+200B), word joiners (U+2060), byte-order marks (U+FEFF), soft hyphens (U+00AD), left-to-right and right-to-left marks, bidirectional embedding, override and isolate controls, and ASCII and C1 control characters. Tabs and line breaks are kept, and Unicode line and paragraph separators become normal line breaks.

Hidden tag characters (U+E0000–U+E007F), which can carry invisible text, are removed too — except inside the flag emoji for England, Scotland and Wales, which need them. Zero-width joiners and non-joiners are kept where they are part of an emoji or of a word in a script that uses them, such as Devanagari or Persian, and removed everywhere else.

Bidirectional controls are the characters behind "Trojan Source" attacks (CVE-2021-42574), which make source code display differently from how it runs.

NFC or NFKC?

NFC (canonical composition) only merges letters and combining accents into single characters. The text looks exactly the same, but searching and comparing become reliable.

NFKC also replaces compatibility characters with plain equivalents: the ligature fi becomes fi, full-width A becomes A, ① becomes 1, superscript ² becomes 2 and the no-break space becomes a normal space. That is useful for search keys, usernames and data cleaning, but it can change meaning — x² becomes x2 — so use NFC for text that people will read.

Limitations

  • Remove HTML tags works on HTML source. It does not apply CSS or run scripts, so text hidden with CSS is kept and text that a page adds with JavaScript is not there to keep.
  • Entity decoding knows every named entity from HTML 4 plus the commonly used HTML5 names; rare HTML5 names such as &NotNestedGreaterGreater; are left unchanged. Numeric references like &#8217; and &#x2019; always work.
  • Emoji are recognised with your browser's Unicode data. Emoji newer than that data may be missed.
  • URL removal recognises links that start with http://, https://, ftp:// or www.. Bare domains such as example.com are kept.
  • Smart-quote conversion cannot tell an apostrophe from a closing single quote; both become '.

Privacy

Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.

Frequently asked questions

Does it run the scripts in my HTML?

No. The HTML is read by a small text parser in this page, not by the browser's HTML engine, so scripts never run and images and other resources are never requested.

How do I find and remove hidden characters?

Keep Remove zero-width and control characters ticked. The count next to it shows how many were found. Text copied from web pages, PDFs and messaging apps often contains zero-width spaces, direction marks or no-break spaces — the last are handled by Normalise Unicode: NFKC or by the Whitespace Remover.

Can I remove line breaks but keep paragraphs?

Yes. Set Line breaks to Join lines, keep paragraphs: lines inside a paragraph are joined with a space, and blank lines between paragraphs stay. Add Remove blank lines to get one paragraph per line.

Why does the result look the same after normalising?

NFC changes how characters are stored, not how they look: an accented letter typed as two code points becomes one. The count shows how many characters were affected.

Is my text uploaded?

No. Every clean-up runs in your browser, and nothing you paste or open is sent to a server.

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.