Remove Punctuation & Characters
Strip punctuation, numbers, emoji or any characters you choose, in one step.
Cleaned in your browser; nothing is uploaded.
Characters
Tick what to remove; everything else stays.
What was removed
| Character | Kind | Times |
|---|
About the Remove Punctuation & Characters
Tick the kinds of characters to remove — punctuation, symbols, numbers, letters, emoji, spaces, line breaks, invisible and control characters or non-ASCII characters — or type your own list, such as #@* or a range like a-z. Switch to Keep only to do the opposite: keep, say, letters and spaces, and remove everything else.
Each kind follows the Unicode character categories, so it works in every language: punctuation includes curly quotes, dashes and the Hindi danda, numbers include ², ½ and Arabic digits, and an emoji with a skin tone or a family emoji is removed as one character. A report shows how many characters each rule removed and which characters were removed most often.
How to use it
- Paste your text or open a text file.
- Choose Remove or Keep only, then tick the kinds of characters. The number after each kind shows how many your text has.
- Optionally type your own characters, keep apostrophes and hyphens inside words, or replace removed characters with a space.
- Copy or download the result, and check What was removed.
Examples
Hello, World!!! Is this the "best"?
Hello World Is this the best
Don’t stop — it’s well-known.
Don’t stop it’s well-known
Punctuation with Keep ’ and - inside words ticked.
Call 98450 12345 or ²½ now
Call or now
Hello, World! 123
Hello World
Hi 👋🏽 family 👨👩👧 ok
Hi family ok
#tag @user *bold*
tag user bold
Your characters: #@*.
Common uses
- Cleaning text for word clouds, word counts, search and machine learning.
- Removing emoji and symbols from names, titles and product data before importing them.
- Finding and removing zero-width spaces, byte order marks and other invisible characters that break code, CSV files and searches.
- Keeping only digits from phone numbers, prices or order IDs.
- Making text ASCII-only for systems that reject other characters.
What each kind covers
Kinds follow the General Category of each character in the Unicode Character Database (UAX #44):
- Punctuation (P): . , ; : ! ? quotes, brackets, dashes, … and the punctuation of other scripts (। ، 。). In Unicode, # % & * @ are punctuation too.
- Symbols (S, without emoji): $ € ₹ + = < > | ~ ^ ` © ® ™ ° arrows and box drawing.
- Emoji: whole emoji sequences, including skin tones, flags, keycaps and ZWJ families.
- Numbers (N): 0–9, digits of other scripts (٣, ३), ², ½ and Roman numerals such as Ⅻ.
- Letters (L): letters of every script, removed together with their accents.
- Spaces and tabs: the normal space, no-break spaces and other space characters, and tabs. Line breaks: LF, CR LF, CR and the Unicode line and paragraph separators.
- Invisible and control characters: format characters such as zero-width spaces and joiners, the byte order mark and soft hyphens, control characters, private-use and unassigned code points. Joiners inside an emoji stay, so the emoji is not broken.
- Non-ASCII: anything outside the 128 ASCII characters, including accented letters and emoji.
Your own characters
Type the characters to match, one after another: #@* matches #, @ and *. A dash between two characters makes a range: a-z, 0-9, A-F. Use \n for a line break, \t for a tab, \s for a space, \- for a dash and \\ for a backslash. Tick Ignore capitals to match a-z and A-Z alike.
Replacing and tidying
Removed characters can be replaced with nothing, a space (so end.Start becomes end Start, not endStart) or your own text. Tidy up spaces then turns double spaces into one and removes spaces at the start and end of each line, which is usually what you want after removing punctuation; untick it to keep spacing exactly.
Limitations
- Works on up to 10 million characters at a time; split longer files into parts with the Text Splitter first.
- Removing zero-width non-joiners changes how some Persian and Indic words are drawn; untick Invisible and control for such text.
- Keeping ’ and - inside words works for apostrophes and hyphens between letters or digits, not for other punctuation.
- Your own characters are matched exactly; regular expressions are not supported (use Find and Replace for patterns).
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
How do I remove all punctuation from text?
Paste the text and keep Punctuation ticked (it is the default). Tick Keep ’ and - inside words to keep contractions like don’t and words like well-known intact.
How do I keep only letters and numbers?
Choose Keep only and tick Letters, Numbers, and Spaces and tabs (and Line breaks to keep lines). Everything else is removed.
How do I remove zero-width spaces and other hidden characters?
Tick Invisible and control. The number next to it shows how many your text has, and What was removed names each one, for example zero-width space U+200B.
Does it work with Hindi, Arabic or Chinese text?
Yes. The kinds follow Unicode categories, so Hindi punctuation (।) counts as punctuation, Devanagari and Arabic digits as numbers, and Hindi, Arabic or Chinese characters as letters.
Is my text uploaded?
No. Everything runs in your browser and nothing you paste leaves your device.