Saved as UTF-8 in Notepad and it still came out garbled
You explicitly chose UTF-8 when saving, and yet opening the file in Excel or another program gives you été or ë³´ê³ ì„œ. Why picking the right option was not enough is what this article is about.
Notepad's save options are confusing
The Save As dialog has an encoding dropdown at the bottom. The entries vary a little by Windows version, but broadly:
- ANSI — sounds like a standard, and is not. More on this below.
- UTF-16 LE — labelled simply "Unicode" in older Notepad. It is Unicode, but it is not UTF-8. Plenty of people pick this and remember it as having saved UTF-8.
- UTF-16 BE — almost never what you want.
- UTF-8
- UTF-8 with BOM — present in some versions, absent in others.
Picking "Unicode" and remembering it as UTF-8 is the single most common version of this mistake. A UTF-16 file opened by something expecting UTF-8 comes out unreadable throughout, or refuses to open at all.
ANSI is not a standard
ANSI is the name of an American standards body, so the entry reads like a specification. In Windows the word actually means "this computer's local code page." That is CP949 on Korean Windows, Shift_JIS on Japanese Windows, Windows-1252 across much of Western Europe.
Which means a file saved as ANSI means different things on different machines. An ANSI file written in Korea garbles on Japanese Windows and vice versa. It happens inside a single company whenever colleagues run different Windows languages.
This is usually why old documents garble now. Back then everyone shared one code page, so nothing looked wrong.
The default changed partway through
Starting with the 2019 update to Windows 10 (version 1903), Notepad's default encoding changed from ANSI to UTF-8. Which means files written before and after that point can sit in the same folder.
So you get the situation where some notes in a folder open fine and others garble. The files genuinely have different encodings, and nothing in the filename or icon distinguishes them.
A BOM is wanted in some places and rejected in others
A BOM is a three-byte marker at the very start of a file. It announces "this is UTF-8" and is invisible on screen. The complication is that some software wants it and some cannot tolerate it.
- Wants it — Excel. Without a BOM it reads a CSV using the local code page. There is more on this in when a CSV comes out garbled in Excel.
- Cannot have it — shell scripts, some config files, older compilers. They read those first three bytes as content, fail to recognise
#!/bin/bash, or throw an error on line one.
So "it is UTF-8" does not settle it. Whether to attach a BOM depends on who receives the file. If your Notepad lists both entries separately, choose according to the destination.
When the file really is UTF-8 and the reader still fails
Sometimes the save was correct and the program opening it simply guesses wrong. Plain text files have nowhere to record their encoding, so without a BOM the reader has to infer it from the content.
The shorter the file, the worse the guess. A note containing a handful of non-Latin characters does not give the detector enough to work with. If the same text opens correctly in a long file but garbles in a short one, this is why.
How to check what encoding a file is in
Open the file in Notepad and the current encoding appears at the right of the status bar along the bottom. If you cannot see it, enable View then Status Bar. Seeing ANSI there rather than UTF-8 means you have found the cause.
That indicator relies on guessing too, though. If the file already looks garbled inside Notepad, treat the reported encoding as suspect as well.
Fixing it with this tool
Drop the file onto the File tab of the Mojibake Recovery tool. It works out how the file was actually saved, shows you the candidates, and you pick the one that reads correctly and download it as UTF-8.
Set the "add a BOM for Excel" checkbox to match your destination when downloading. It is on by default for .csv, .tsv and .txt. Turn it off for shell scripts and config files.
Files are never uploaded. Everything happens inside the browser, so work documents are fine to use as they are.
Avoiding it next time
- Save anything you are sending to someone else as UTF-8. ANSI changes meaning with the recipient's system language.
- Be wary of an entry labelled "Unicode." It usually means UTF-16, not UTF-8.
- Add a BOM for CSVs headed to Excel, and for most other purposes leave it off.
- If you want to understand why any of this happens, see UTF-8 and legacy encodings, explained. Knowing the cause lets you reason about the next situation too.