Context
Task 1a gave me 248 Abt and 249 Buy product listings (the course's subset of the public Abt-Buy benchmark) and asked for the pairs that describe the same product. A staff script scored the output by precision and recall against 149 true matches. Both shops write the same product differently: Abt names usually end with a model code ("Sony SLV-D380P Black DVD VHS Combo Player - SLVD380P"), Buy names often contain the code with punctuation removed ("Sony SLVD380P DVD/VCR Combo"), descriptions are long and noisy, and the Buy manufacturer field is often blank. There was no labelled training split to learn from, only the data and the scorer.
Decision
Link records in four passes of decreasing strictness, where each pass only considers records the earlier passes left unmatched.
- Model number: the last word of the Abt name appears verbatim in the Buy
name (with
-and/removed). - Brand and model: the same first word (brand), and the Abt model code minus its last character appears in the Buy name.
- Description similarity within a brand: the mean of token overlap, token
Jaccard and
fuzz.ratiobetween descriptions; keep the best candidate if it clears a threshold (0.4 in the report, 0.5 in the submitted script). - Name similarity:
fuzz.ratiobetween names, kept above 0.325.
Options considered
- One similarity score over every pair with a single threshold. Simple, but one threshold cannot serve both the exact model-code cases and the vague description-only cases.
- A trained classifier on pair features. The usual answer, but there was no training split, and fitting on the scored pairs would have meant tuning to the test.
- Model numbers only (passes 1 and 2). Very precise, but it misses every product whose listing has no clean model code.
- The four-pass cascade (chosen). Spend the strong evidence first, then fall back to fuzzier evidence only for what is left.
Why
Model codes are the strongest and most precise signal in these catalogues, so they should decide first. Restricting each later pass to unmatched records keeps the output close to one-to-one and stops the fuzzy passes from overriding confident matches. Grouping pass 3 by brand cuts the number of description comparisons and removes a large class of false positives.
What happened
- As reported: precision 0.8639, recall 0.8523 (127 correct of 147 pairs output, 149 true matches). The revival adds 95% Wilson intervals of [0.799, 0.910] for precision and [0.787, 0.900] for recall, and an F1 of 0.858 with a bootstrap interval of [0.810, 0.902] (resampling Abt records, seed 2020).
- The passes are very uneven. Pass 1 found 93 pairs (89 correct) and pass 2 found 27 (all correct). Pass 3 found 19, only 10 correct, and pass 4 found 8 with just 1 correct: on its own, the last pass made the result worse.
- Re-running the code in 2026 showed that the report's numbers come from a
pass-3 threshold of 0.4 with fuzzywuzzy's pure-Python backend, while the
submitted script says 0.5 and its saved
task1a.csvneeds the python-Levenshtein backend. The report did not say this. - Three bugs change the output and are kept in the port for fidelity: pass 3 marks the last Buy record as used instead of the one it matched (a leaked loop variable), pass 4 compares a 0 to 100 score with a running maximum stored on a 0 to 1 scale (so the last positive candidate wins), and pass 4 only scans the first n rows of the Buy table, where n is the number of unmatched Buy records.
- In 2026 the LLM evaluation harness on the site (
/llm-eval) puts this matcher, exactly as reported, against an LLM judge on a sample of pairs (see DR-004).
What I'd change
- Drop pass 4, or replace it with a check that only fires when the name similarity is high and the brands agree.
- Choose thresholds on a validation split, not by eye against the scorer, and report intervals from the start.
- Replace the greedy pass order with a global one-to-one assignment on a combined score, which removes the "used record" bookkeeping where the bugs lived.
- Keep the code and the report in sync: the 0.4 versus 0.5 mismatch is the kind of thing a single config file and a results table generated by the code would have prevented.