Text turns into gibberish like “é”, “’” or “Привет” when it was saved in one encoding and opened in another. The letters are not gone: each one has been split into the wrong characters. This fixer works out which two encodings were mixed up, by trying every likely pair and keeping the reading that looks most like real writing, and turns the text back. It handles UTF-8 read as Windows-1252, Latin-1, Mac Roman or the old DOS code page, text garbled twice over, Cyrillic, Greek, Central European, Turkish, Hebrew and Arabic code pages, and Chinese, Japanese and Korean text read as the wrong one. Text that is only partly garbled is fixed without touching the rest.
It also opens text files whose encoding you do not know, such as an old .txt, a .csv that Excel shows with odd characters, or subtitles from a decade ago. It works out the encoding, shows a preview, and saves the file as UTF-8, which every modern program reads. It is free to use, with no account and no upload limit, and it runs entirely in your browser: nothing is uploaded, and while the page stays open, a feature you have already used keeps working offline.
- What is mojibake?
- It is the Japanese word for garbled characters. It happens when text is written in one encoding and read in another. Most often, UTF-8 is read as Windows-1252, so an “é” (two bytes in UTF-8) is shown as the two characters “é” and a curly apostrophe becomes “’”. Nothing is lost: the bytes are all still there, read the wrong way, so reading them the right way brings the text back.
- Why can some characters not be fixed?
- When text shows a black diamond with a question mark (�), some program along the way met bytes it could not read and replaced them with that symbol. The original bytes were thrown away at that point, so no tool can recover them from the text. Go back to the original file if you have it: open it here and the encoding can be read correctly from the start.
- Why does my CSV show strange characters in Excel?
- Excel assumes a CSV without a byte order mark is in your computer’s old code page, so a UTF-8 file comes out garbled. Open the CSV here, keep “Add a byte order mark” ticked, and save it. Excel then reads it as UTF-8 and the accents and symbols appear correctly.
- How does it know which encoding a file is in?
- A byte order mark at the start settles it, and so does text that is valid UTF-8, because older encodings almost never are by accident. Otherwise each code page is tried on the start of the file, and the one that gives the most natural reading wins. That is a well-informed guess, which is why the preview is there and other encodings are a click away.
- Can it save in encodings other than UTF-8?
- No. Files are saved as UTF-8, which holds every character of every language and is what modern software expects. If an old program needs a legacy code page, convert the UTF-8 file with that program or with a tool such as iconv.