Skip to content
Data Processing LabUniMelb · 2020 S2

Linkage · LLM evaluation · added in 2026

LLM judge vs the 2020 matcher

Would a language model link Abt and Buy products better than the four hand-written passes from 2020? This harness asks both the same question about the same pairs and scores them against the ground truth, with intervals instead of single numbers. The matcher runs in your browser; the LLM runs only if you add your own API key.

candidate pairs
3,054candidate pairs
of them matches
4.9%of them matches
pairs per run (default)
40pairs per run (default)
sample seed
2020sample seed
Evaluation design

Same pairs, not the same view

Both methods are scored on the same pairs against the same ground truth, so every comparison is paired. They do not get the same information, though: the matcher works on the whole catalogue, the LLM on one pair at a time, and the results say how much that matters. The design choices, and why, are in decision record DR-004.

The pairs

Candidate pairs are what blocking would hand a matcher: every Abt × Buy pair that shares the submitted brand key, plus every true match (3,054 pairs, 4.9% matches). A run samples half matches and half same-brand non-matches with a fixed seed, never looking at either method's answers, and never uses a listing in two pairs, so the pairs are independent.

The contestants

The 2020 matcher exactly as reported (pass 3 at 0.4, pure-Python difflib), run on all 248 × 249 records; a pair counts as matched if it is among its output. That run has context the LLM lacks: a listing already linked is skipped, and passes 3 and 4 keep only the best candidate. The LLM sees only the two listings (names, descriptions cut to 500 characters, Buy manufacturer) and returns a verdict, a confidence and a one-line rationale as validated JSON.

The scoring

Precision, recall, specificity and accuracy with Wilson intervals; F1 and LLM-minus-matcher differences with stratified, paired bootstrap intervals; McNemar's exact test on the pairs where only one method was right; precision re-weighted to the pool's real match rate; each score split by whether the matcher had linked a listing elsewhere; tokens, latency and estimated cost.

What this cannot tell you: how either method does on catalogues other than this course's 248 × 249 subset, or how a judge does with a different prompt. Forty pairs give wide intervals; the point is to see how wide, not to crown a winner.

Bring your own key

Run the judge

Pick a sample, then run the judge on your own Anthropic or OpenAI key. The calls go from your browser straight to the provider; each one is written to the AI log in this browser. Without a key, everything below still works for the matcher.

Running the 2020 matcher on all 248 × 249 records in a background worker…
2020 matcher

Results

Scores appear when the matcher finishes.
Human in the loop

Read the disagreements

The sampled pairs appear when the matcher finishes.
Transparency

How the AI part is governed

The design is informed by the Australian Government's policy for the responsible use of AI in government, the EU AI Act's transparency principles and the NIST AI Risk Management Framework. It is a portfolio demonstration, not a certified or audited system.

In practice

  • Nothing calls a model until you add your own key and press Run.
  • Your key stays in your browser and goes only to the provider you chose; this site has no server.
  • Every AI answer is labelled AI-generated where it appears.
  • Every call that leaves your browser (prompt, parameters, answer or error, latency, tokens, your review) is logged in this browser and can be exported, including calls you stop mid-flight.
  • The scores use the model's own answers; your accept, flip or reject decisions are recorded beside them, never silently merged in.

Read more

The AI use statement says what the AI does here, what it never does and what data goes to the provider. The AI log shows every call made from this browser.

Data source: Abt-Buy product matching benchmark, Leipzig University database group; the subset used in COMP20008; ids, names, descriptions and manufacturers only.