Skip to content
Data Processing LabUniMelb · 2020 S2

Decision record · DR-001

A four-pass matcher for Abt-Buy record linkage

Link records in four passes of decreasing strictness, where each pass only considers records the earlier passes left unmatched.

Status
Accepted in 2020, recorded retrospectively in October 2026
Decided
2020 Semester 2, Project 2 task 1a (coursework/project-2/task1a.py)
Recorded
2026-10-06, while reviving the coursework

Context

Task 1a gave me 248 Abt and 249 Buy product listings (the course's subset of the public Abt-Buy benchmark) and asked for the pairs that describe the same product. A staff script scored the output by precision and recall against 149 true matches. Both shops write the same product differently: Abt names usually end with a model code ("Sony SLV-D380P Black DVD VHS Combo Player - SLVD380P"), Buy names often contain the code with punctuation removed ("Sony SLVD380P DVD/VCR Combo"), descriptions are long and noisy, and the Buy manufacturer field is often blank. There was no labelled training split to learn from, only the data and the scorer.

Decision

Link records in four passes of decreasing strictness, where each pass only considers records the earlier passes left unmatched.

  1. Model number: the last word of the Abt name appears verbatim in the Buy name (with - and / removed).
  2. Brand and model: the same first word (brand), and the Abt model code minus its last character appears in the Buy name.
  3. Description similarity within a brand: the mean of token overlap, token Jaccard and fuzz.ratio between descriptions; keep the best candidate if it clears a threshold (0.4 in the report, 0.5 in the submitted script).
  4. Name similarity: fuzz.ratio between names, kept above 0.325.

Options considered

  • One similarity score over every pair with a single threshold. Simple, but one threshold cannot serve both the exact model-code cases and the vague description-only cases.
  • A trained classifier on pair features. The usual answer, but there was no training split, and fitting on the scored pairs would have meant tuning to the test.
  • Model numbers only (passes 1 and 2). Very precise, but it misses every product whose listing has no clean model code.
  • The four-pass cascade (chosen). Spend the strong evidence first, then fall back to fuzzier evidence only for what is left.

Why

Model codes are the strongest and most precise signal in these catalogues, so they should decide first. Restricting each later pass to unmatched records keeps the output close to one-to-one and stops the fuzzy passes from overriding confident matches. Grouping pass 3 by brand cuts the number of description comparisons and removes a large class of false positives.

What happened

  • As reported: precision 0.8639, recall 0.8523 (127 correct of 147 pairs output, 149 true matches). The revival adds 95% Wilson intervals of [0.799, 0.910] for precision and [0.787, 0.900] for recall, and an F1 of 0.858 with a bootstrap interval of [0.810, 0.902] (resampling Abt records, seed 2020).
  • The passes are very uneven. Pass 1 found 93 pairs (89 correct) and pass 2 found 27 (all correct). Pass 3 found 19, only 10 correct, and pass 4 found 8 with just 1 correct: on its own, the last pass made the result worse.
  • Re-running the code in 2026 showed that the report's numbers come from a pass-3 threshold of 0.4 with fuzzywuzzy's pure-Python backend, while the submitted script says 0.5 and its saved task1a.csv needs the python-Levenshtein backend. The report did not say this.
  • Three bugs change the output and are kept in the port for fidelity: pass 3 marks the last Buy record as used instead of the one it matched (a leaked loop variable), pass 4 compares a 0 to 100 score with a running maximum stored on a 0 to 1 scale (so the last positive candidate wins), and pass 4 only scans the first n rows of the Buy table, where n is the number of unmatched Buy records.
  • In 2026 the LLM evaluation harness on the site (/llm-eval) puts this matcher, exactly as reported, against an LLM judge on a sample of pairs (see DR-004).

What I'd change

  • Drop pass 4, or replace it with a check that only fires when the name similarity is high and the brands agree.
  • Choose thresholds on a validation split, not by eye against the scorer, and report intervals from the start.
  • Replace the greedy pass order with a global one-to-one assignment on a combined score, which removes the "used record" bookkeeping where the bugs lived.
  • Keep the code and the report in sync: the 0.4 versus 0.5 mismatch is the kind of thing a single config file and a results table generated by the code would have prevented.