The two invented versions
This is a synthetic worked example with invented values, not a customer case or a supplier's file. Results marked computed were produced with Python's standard library on invented input. Version 2 is re-sorted, so a line comparison would flag nearly every line. The key is the sku column.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
OLD (8 rows) NEW (9 rows, re-sorted)
A100,Hinge brass,4.20,120 A107,Plug,0.90,300
A101,Bolt M6,1.10,500 A100,Hinge brass,4.2,120
A102,Washer,0.05,2000 A101,Bolt M6,1.65,500
A103,Nut M6,0.30,900 A102,Washer,0.05,0
A104,Screw,0.12,0 A103,Nut M6,0.30,900
A105,Bracket,2.50,40 A105,Bracket,2.50,40
A106,Clamp,3.80,75 A106,"Clamp, steel",3.80,75
A107,Plug,0.90,300 A108,Hex key,1.40,60
A109,Spanner,6.00,25The report a key-based comparison should give
After normalising (trim and collapse spaces, compare prices and stock as numbers) the report below is what the comparison should produce: six real differences, three changed products, two added and one absent. Two differences are formatting only and must not appear: A100's price written 4.2 instead of 4.20, and A103's name with a doubled space. Computed with Python's csv and decimal modules.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
ADDED (2): A108 Hex key; A109 Spanner
ABSENT from new file (1): A104 Screw (absent, not 'discontinued')
CHANGED (3):
A101 price 1.10 -> 1.65 (+50%) over the 20% threshold
A102 stock 2000 -> 0 stock fell to zero
A106 name Clamp -> Clamp, steel
NOT REPORTED (formatting only): A100 price '4.20' -> '4.2'; A103 name spacing; row orderThe counts check
The reconciliation is on distinct keys. The old file has 8 distinct keys, minus 1 absent, plus 2 added, equals 9. The new file has 9 distinct keys, so the report reconciles. Had the new file shown 10, a key would have been lost or double counted and the report would be wrong. No key repeats in either file here, so the distinct keys and the rows are the same numbers, and the extra rows that repeat a key are 0 in each file. Comparing the file with itself must give no changes, and running it twice must give byte-identical reports.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
old file new file
rows 8 9
extra rows repeating a key 0 0
distinct keys 8 9
old distinct keys 8
absent from new -1
added in new +2
expected new distinct keys 9
actual new distinct keys 9 -> reconcilesWhat this example does not show
It is not a real supplier's change history and it says nothing about what supplier prices should be. A key that repeats in either file would be listed with all its rows, not compared and not resolved, and the extra rows that repeat a key would be counted separately for each file so that the distinct-key reconciliation still balances; this example has no repeated key. The size is tiny; the job's acceptance test uses ten planted differences, including a key that repeats in only one file. No data was imported or changed anywhere.
Use it to specify a priced enquiry
The fixed job feed-version-diff-report-before-import is £225 for one supplier layout and two file versions of up to 100,000 rows each; its acceptance is a planted-change test like this one, a zero-change self-comparison and a reconciliation on your real pair. Send column names, row counts and the thresholds you care about first, never real price lists. Prices are untested proposals, and payment follows the agreed checks and your sign-off. Nothing is booked or charged by an enquiry.
Sources and limits
- Python difflib documentation Checked 2026-10-11.
- difflib compares sequences, usually of text lines, and its default junk heuristic is asymmetric.
- Python csv documentation Checked 2026-10-11.
- DictReader reads rows into dicts keyed by header names.