Skip to content
Data Processing LabUniMelb · 2020 S2

Decision record · DR-004

How to evaluate an LLM judge against the 2020 matcher

Evaluate the matcher and an LLM judge as a paired comparison on a seeded, stratified sample of candidate pairs, with every LLM call made on the visitor's own key and written to an audit log.

Status
Accepted
Decided
2026-10, when adding the LLM evaluation harness (/llm-eval)
Recorded
2026-10-06

Context

The obvious 2026 question about task 1a is whether a large language model would link the two catalogues better than four hand-written passes. Answering it honestly needs the same pairs, the same ground truth and uncertainty around every number. There were also two hard constraints: no budget for API calls, and a static site with no server. Any AI feature therefore has to run on the visitor's own key, directly from their browser, and the site has to work fully without one.

Decision

Evaluate the matcher and an LLM judge as a paired comparison on a seeded, stratified sample of candidate pairs, with every LLM call made on the visitor's own key and written to an audit log.

  • Pool: every Abt × Buy pair that shares the submitted brand blocking key, plus every true match (3,054 pairs, 149 matches, a 4.9% match rate).
  • Sample: half matches and half same-brand non-matches (default 40, seed 2020), drawn without looking at either method's answers, and record-disjoint: no Abt or Buy listing appears in two sampled pairs, so the pairs are independent units for the intervals and the test.
  • Matcher: the 2020 code exactly as reported (pass 3 at 0.4, pure-Python backend) run on all 248 × 249 records; a pair counts as matched if it is in its output. Its verdicts therefore use catalogue-wide context the judge does not get: passes 2 to 4 skip any listing already linked, and passes 3 and 4 keep only the best candidate for each Abt listing.
  • Judge: sees only the two listings (names, descriptions cut to 500 characters, the Buy manufacturer), returns a verdict, a confidence and a one-line rationale as JSON validated against a schema. Claude Haiku 4.5 by default (temperature 0), Claude Sonnet 5.5 at low effort as an option, or an OpenAI model the visitor names.
  • Scoring: Wilson intervals for precision, recall, specificity and accuracy; stratified, paired percentile bootstrap for F1 and for the LLM-minus-matcher differences; McNemar's exact test on per-pair correctness; precision re-weighted to the pool's 4.9% match rate; and specificity and recall for both methods split by whether the matcher had linked one of the pair's listings elsewhere, so the difference in what each method sees is measured rather than hidden.
  • No silent model switching: no server-side fallback to another model, so the model that answered is always the model that is recorded.

Options considered

  • Random pairs from all 61,752. Almost all are obvious non-matches, so both methods would look near perfect and the comparison would say nothing.
  • Sampling from either method's output (for example the matcher's mistakes). Selecting on the outcome biases the comparison towards the method that did not choose the sample.
  • Judging the whole pool. Fair and precise, but about 3,000 calls would cost a visitor a few dollars on the cheapest model. The sample size is a choice on the page instead (20 to 120).
  • A server proxy with my own key. No budget, and it would mean holding a key on a server for a portfolio site.
  • Percentile bootstrap for every metric. Tried first; the unit tests showed it collapses to a single point when a method makes no false positives in the sample (the matcher on the default sample: precision interval [1, 1]), so proportions now use Wilson intervals.
  • Resampling by record instead of by pair (a cluster bootstrap), keeping a sampler that lets pairs share listings. It would fix the bootstrap but not the Wilson intervals or McNemar's test, which also assume independent pairs; a record-disjoint sample fixes all three.

Why

Pairing removes the variation between pairs from the comparison, which is what makes 40 pairs usable at all. Stratifying guarantees enough matches to estimate recall; same-brand negatives are the cases blocking would actually hand a matcher. Re-weighting precision to the pool's match rate stops a balanced sample from flattering a method that says "match" too readily. Bring-your-own-key keeps the cost with the person choosing to run it, and the audit log makes every AI output inspectable afterwards.

What happened

  • The harness, the statistics and the provider adapters are unit-tested with mocked API responses (no real calls in tests or in CI).
  • Review before merging found two flaws in the first version, both fixed before anything was published:
    • The sampler drew non-matches from the same brands as the sampled matches and often reused their listings: 11 of the 40 default pairs shared a listing with another sampled pair (about 10 of 40 on average across seeds, about 63 of 120 at the largest size), while the methods page claimed they "rarely" did. The sampler now skips any pair that would reuse a listing, and a test checks every offered size over 25 seeds.
    • The page called the comparison "a fair fight" without saying that the matcher sees the whole catalogue. In the default sample 15 of the 20 non-matches involve a listing the matcher had linked to something else; across the pool it is 2,379 of 2,905 non-matches. The matcher's specificity on the pool is high either way (0.997 with that context, 0.990 without), but the judge can only ever use the second kind of evidence. The page now says so and reports the split.
  • The Anthropic adapter first built its schema with the SDK's zod helper, which (in SDK 0.131) moves string enums into the description, so the API only enforced "a string" for the verdict and confidence. Both adapters now send the same schema with the enums intact, and a test checks them.
  • No LLM results are published on the site. There was no budget to run the judge, and publishing numbers I could not reproduce for a visitor would defeat the purpose. Visitors who run it see their own results and can export them.
  • The matcher alone on the default sample of 40 (the record-disjoint sample; its counts happen to equal the first version's): 39 correct (precision 1.000 [0.832, 1.000], recall 0.950 [0.764, 0.991]). Its precision re-weighted to the pool's match rate is 1.000 with an interval of [0.157, 1.000]: with only 20 non-matches, a sample cannot pin down precision at a 4.9% match rate. Its real precision on the pool, known from the full run, is 0.914 (127 of 139 predicted pairs), inside that interval. That width is the most useful thing the harness shows.

What I'd change

  • Sample many more non-matches than matches (or weight them by inverse inclusion probability), because precision at a low match rate depends almost entirely on the false-positive rate.
  • Fix the judge prompt on one sample and evaluate it on a fresh one, so prompt tweaks cannot overfit the evaluation pairs.
  • Check whether the judge's stated confidence is calibrated before using it to route pairs.
  • Try the judge as a reviewer of the matcher's weakest passes (3 and 4) rather than as a replacement for the whole matcher.
  • Add a like-for-like baseline: the four passes applied to one pair at a time, without the one-to-one exclusion or the best-candidate step. Or give the judge the same view as the matcher, asking it to pick the best Buy listing from a brand block.