Text Encoding Converter & Mojibake Fixer
Read old text files correctly, save them as UTF-8, and un-garble “café”.
If letters look wrong in the preview, choose another encoding — the preview updates at once.
Compare the likeliest encodings
This text looks like mojibake: UTF-8 that was read with the wrong code page at some point (for example “café” instead of “café”).
Save as
Single-byte code pages hold at most 256 characters.
| Character | Code point | Count | First at | Written as |
|---|
Repaired on your device as you type — nothing is uploaded.
Automatic mode only changes text that is clearly garbled. Choose an encoding to force a repair.
What changed
| Garbled | Repaired | Times |
|---|
About the Text Encoding Converter & Mojibake Fixer
A text file is just bytes; the encoding says which bytes mean which letters. Open a file with the wrong one and you get gibberish — accents turn into é, Russian into при, Japanese into boxes. This tool reads the bytes of your file, works out the most likely encoding (byte-order mark, UTF-8 validity, UTF-16 zero-byte patterns, then a statistical detector), shows a live preview, and lets you switch encodings until the text looks right. Then it saves the file as UTF-8 (with or without a BOM), UTF-16, or back to a single-byte code page such as Windows-1252.
The second mode repairs mojibake: text whose UTF-8 bytes were already decoded with the wrong code page — often more than once — so café becomes café and it’s becomes it’s. Everything runs in your browser with its built-in decoders from the WHATWG Encoding Standard; your file is never uploaded.
How to use it
- Choose Convert a file and open or drop the file. The detected encoding is selected for you and the preview shows the text.
- If letters look wrong, pick another encoding under Read the file as, or press Use next to one of the likeliest encodings — each shows its version of the first line with accented letters.
- If the preview contains sequences like
éor’, press Repair it: the text was UTF-8 that someone read with the wrong code page. - Under Save as, choose UTF-8 (recommended), UTF-8 with BOM, UTF-16 or a code page, and the line endings you need. For a code page, choose what happens to characters it does not have.
- Press Download converted file. To fix pasted text instead of a file, switch to Fix garbled text (mojibake) and paste it.
Examples
café crème, Größe
café crème, Größe
it’s 5 € 😀
it’s 5 € 😀
café
café
Repaired in two passes; the change list shows é → é.
Привет
Привет
Łódź
Lódz
Windows-1252 has “ó” but not “Ł” or “ź”; with Stop and list them the tool reports both instead.
Common uses
- Converting old Windows “ANSI” (Windows-1252, Windows-1251, Windows-1250) subtitle, CSV and text files to UTF-8 for modern apps and websites.
- Reading Japanese (Shift_JIS, EUC-JP), Chinese (GBK, GB18030, Big5) and Korean (EUC-KR) files that show as gibberish.
- Making a CSV open correctly in Excel by saving it as UTF-8 with a BOM.
- Fixing product descriptions, emails and database exports full of
é,’andÂafter a bad import. - Saving text back to a legacy code page for an old program, and finding exactly which characters it cannot store.
How the encoding is detected
- Byte-order mark.
EF BB BFmeans UTF-8,FF FEUTF-16 LE,FE FFUTF-16 BE; UTF-32 marks are recognised too. - Valid UTF-8. If every byte sequence in the file is valid UTF-8, it almost certainly is UTF-8: legacy text with accents practically never forms valid UTF-8 by accident. Pure ASCII files read the same in UTF-8 and every Western code page.
- Zero bytes. UTF-16 text without a BOM has its zero bytes on one side of each byte pair — in every other byte for English, at the spaces, digits and line breaks for Chinese or Russian — and decodes to ordinary characters; other zero bytes suggest a binary file.
- Statistics. Otherwise the open-source detector chardet compares letter and byte patterns with typical text in each language and encoding. Each guess is then decoded, and guesses that produce undecodable bytes or control characters are moved down.
Short files give the statistics little to work with, so always check the preview — the comparison list makes the right choice obvious in most cases.
What mojibake is and how the repair works
UTF-8 stores “é” as the two bytes C3 A9. A program that wrongly assumes Windows-1252 shows those two bytes as the two characters à and ©. If that garbled text is saved as UTF-8 and read wrongly again, it doubles: é.
The repair runs this backwards: it turns the garbled characters back into the bytes the wrong code page gave them and decodes those bytes as UTF-8. It tries Windows-1252 (and its ISO-8859-1 variant), Windows-1250, 1251, 1253, 1254, 1257, ISO-8859-2, ISO-8859-15 and Mac OS Roman, repeats up to three times for double encoding, and lists every change. In automatic mode a repair is only applied when the repaired sequences make up at least half of the text’s non-ASCII characters, so correctly written accented text is normally left alone. A word or two gives little evidence — the Russian word “Её” is byte for byte what a garbled “Ÿ” looks like — so check the list of changes for very short text.
UTF-8, BOM or a code page?
UTF-8 stores every character of every language and is the standard for the web, Linux, macOS and modern Windows apps — choose it unless something else is required.
UTF-8 with BOM adds the three bytes EF BB BF at the start. Excel on Windows needs them to open a UTF-8 CSV correctly, and some old Windows programs use them to recognise UTF-8. Leave the BOM out for web pages, scripts, JSON and configuration files, where it can show up as stray characters.
UTF-16 is what Windows calls “Unicode” in Notepad and PowerShell.
Code pages such as Windows-1252 hold at most 256 characters. Use them only when an old program or device insists; the tool tells you which characters do not fit.
Limitations
- Saving is possible as UTF-8, UTF-16 and the 28 single-byte code pages listed. Browsers only include decoders for Shift_JIS, EUC-JP, ISO-2022-JP, GBK, GB18030, Big5 and EUC-KR, so files in those encodings can be read and converted to UTF-8, but not written back.
- DOS code pages 437 and 850 and EBCDIC are not part of the WHATWG Encoding Standard, so browsers cannot read them. IBM866 (DOS Cyrillic) is supported.
- UTF-32 files can be read but are saved as UTF-8 or UTF-16.
- Detection is a well-informed guess, especially for short files and for code pages that share most letters (for example Windows-1250 and ISO-8859-2). The preview is the final judge.
- Characters already shown as � (U+FFFD) were lost when the text was decoded and cannot be restored from the text. Open the original file instead and choose the right encoding.
- The repair handles UTF-8 read as the single-byte code pages listed above. UTF-8 read as Shift_JIS or GBK (for example “譁蟄怜喧縺�”) cannot be undone without encoders the browser does not have.
- Files up to 50 MB.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
Why do I see é, ’ or  in my text?
The text was saved as UTF-8 but read as Windows-1252 (or ISO-8859-1). Each accented letter or curly quote is two or three bytes in UTF-8, and the wrong code page shows each byte as a separate character. Paste the text into Fix garbled text or open the file and press Repair it.
How do I convert an ANSI file to UTF-8?
“ANSI” in Windows means the system code page — Windows-1252 for Western European languages, Windows-1251 for Cyrillic, Windows-1250 for Central European. Open the file, check that the preview reads correctly (switch encodings if not), keep UTF-8 under Save as and download.
Should I save with a BOM?
Only if the file is a CSV for Excel or is used by an old Windows program that needs it. For web pages, code, JSON and anything used on Linux or macOS, save UTF-8 without a BOM.
What is the difference between ISO-8859-1 and Windows-1252?
Windows-1252 is ISO-8859-1 plus printable characters in the range 0x80–0x9F, such as €, “ ”, ‘ ’, – and —. Browsers follow the WHATWG Encoding Standard and read files labelled ISO-8859-1 as Windows-1252, so this tool lists them together.
Why does the tool refuse to save my text as Windows-1252?
The text contains characters that code page does not have, such as emoji, “Ł” or Chinese characters. The table lists each one with its first position. Save as UTF-8, or choose closest ASCII, ? or HTML numeric reference for the missing characters.
Is my file uploaded?
No. The file is read and converted by your browser on your device; nothing is sent to a server, and the page keeps working offline once loaded.