Linkage · LLM evaluation · added in 2026
LLM judge vs the 2020 matcher
Would a language model link Abt and Buy products better than the four hand-written passes from 2020? This harness asks both the same question about the same pairs and scores them against the ground truth, with intervals instead of single numbers. The matcher runs in your browser; the LLM runs only if you add your own API key.
- candidate pairs
- 3,054candidate pairs
- of them matches
- 4.9%of them matches
- pairs per run (default)
- 40pairs per run (default)
- sample seed
- 2020sample seed
Same pairs, not the same view
Both methods are scored on the same pairs against the same ground truth, so every comparison is paired. They do not get the same information, though: the matcher works on the whole catalogue, the LLM on one pair at a time, and the results say how much that matters. The design choices, and why, are in decision record DR-004.
The pairs
Candidate pairs are what blocking would hand a matcher: every Abt × Buy pair that shares the submitted brand key, plus every true match (3,054 pairs, 4.9% matches). A run samples half matches and half same-brand non-matches with a fixed seed, never looking at either method's answers, and never uses a listing in two pairs, so the pairs are independent.
The contestants
The 2020 matcher exactly as reported (pass 3 at 0.4, pure-Python difflib), run on all 248 × 249 records; a pair counts as matched if it is among its output. That run has context the LLM lacks: a listing already linked is skipped, and passes 3 and 4 keep only the best candidate. The LLM sees only the two listings (names, descriptions cut to 500 characters, Buy manufacturer) and returns a verdict, a confidence and a one-line rationale as validated JSON.
The scoring
Precision, recall, specificity and accuracy with Wilson intervals; F1 and LLM-minus-matcher differences with stratified, paired bootstrap intervals; McNemar's exact test on the pairs where only one method was right; precision re-weighted to the pool's real match rate; each score split by whether the matcher had linked a listing elsewhere; tokens, latency and estimated cost.
What this cannot tell you: how either method does on catalogues other than this course's 248 × 249 subset, or how a judge does with a different prompt. Forty pairs give wide intervals; the point is to see how wide, not to crown a winner.
Run the judge
Pick a sample, then run the judge on your own Anthropic or OpenAI key. The calls go from your browser straight to the provider; each one is written to the AI log in this browser. Without a key, everything below still works for the matcher.
Results
Read the disagreements
How the AI part is governed
The design is informed by the Australian Government's policy for the responsible use of AI in government, the EU AI Act's transparency principles and the NIST AI Risk Management Framework. It is a portfolio demonstration, not a certified or audited system.
In practice
- Nothing calls a model until you add your own key and press Run.
- Your key stays in your browser and goes only to the provider you chose; this site has no server.
- Every AI answer is labelled AI-generated where it appears.
- Every call that leaves your browser (prompt, parameters, answer or error, latency, tokens, your review) is logged in this browser and can be exported, including calls you stop mid-flight.
- The scores use the model's own answers; your accept, flip or reject decisions are recorded beside them, never silently merged in.
Read more
The AI use statement says what the AI does here, what it never does and what data goes to the provider. The AI log shows every call made from this browser.
Data source: Abt-Buy product matching benchmark, Leipzig University database group; the subset used in COMP20008; ids, names, descriptions and manufacturers only.