The invented file
This is a synthetic worked example with invented values, not a customer case or a supplier's file. Results marked computed were produced with Python's standard library on invented input, and each reading below names its encoding explicitly. The file below is shown as text; its first three bytes, EF BB BF, are a byte-order mark that no editor displays. The product names and prices are made up.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
bytes 1-3: EF BB BF (byte-order mark, invisible)
line 1: sku,name,price
line 2: A100,Café table,129.00
line 3: A101,Müller® tray – large,19.50
line 4: A102,Euro mat (€ 4.99) ™,4.99Two readings of the first header
Read as plain UTF-8, the first header is the four characters U+FEFF, s, k and u. A lookup for sku finds nothing, so the importer reports a missing column or skips the first column. Read with the utf-8-sig codec, the three bytes are dropped and the header is sku. Computed with Python's csv module and codecs.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
plain utf-8 : first header = U+FEFF + 'sku' -> equals 'sku'? False
utf-8-sig : first header = 'sku' -> equals 'sku'? TrueWhat the wrong reading of the accents looks like
If the same bytes are read as the Windows single-byte code page instead of UTF-8, each accented character becomes two or three characters. The table is computed, not copied from a supplier.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
character | UTF-8 bytes | read as Windows-1252
é | C3 A9 | é
ü | C3 BC | ü
– | E2 80 93 | –
€ | E2 82 AC | €
™ | E2 84 A2 | â„¢The assertions an importer test should make
Use the same file and assert these results. The stored name for A100 equals the text Café table with the accented letter intact. The stored name for A101 contains the umlaut, the registered sign and the en dash. The stored name for A102 contains the euro sign and the trademark sign. The first header equals sku. The file saved without the byte-order mark gives identical stored values. A copy with an invalid byte, for example a lone byte that does not form a valid UTF-8 sequence, is rejected with a message that names the line, and none of its rows is stored. The strict handling matters: Python's documentation says the replace handler substitutes U+FFFD and ignore drops bad bytes silently, so either would store damage.
- Compare as code points, not by eye; look-alike characters hide differences.
- Run the same assertions on every supplier layout the importer reads.
What this example does not show
It does not show how any particular supplier writes its files, which spreadsheet program adds a byte-order mark, or whether your importer already handles these cases. It does not say which encoding any particular language or machine assumes when none is named: Python's csv documentation describes the default differently for 3.14 (the system default encoding) and 3.15 (UTF-8), so an importer test should name the encoding it expects, as the guide on encodings and the byte-order mark explains. It does not repair text that is already stored garbled, and it cannot recover characters a previous save destroyed. It is an authored fixture, not an observation, and no data was imported anywhere.
Use it to specify a priced enquiry
If your importer fails one of these assertions, the fixed job csv-encoding-bom-garbled-characters-import is £195 for one importer path and one named supplier layout. Send three to ten invented or redacted rows and the importer's name first, never real price lists, credentials or code. Prices are untested proposals, and payment follows the agreed checks and your sign-off. Nothing is booked or charged by an enquiry.
Sources and limits
- Python codecs documentation Checked 2026-10-11.
- The utf-8-sig codec skips a leading byte-order mark when decoding, while plain utf-8 keeps U+FEFF as an ordinary character.
- Python 3.14 csv documentation Checked 2026-10-11.
- Because open() is used to read a CSV file, the file is by default decoded using the system default encoding (see locale.getencoding()); another encoding is chosen by passing the encoding argument of open().
- Python 3.15 csv documentation Checked 2026-10-11.
- Because open() is used to read a CSV file, the file is by default decoded using UTF-8, so the documented default differs from Python 3.14.