Skip to content
Data Processing LabUniMelb · 2020 S2

Decision record · DR-002

Blocking the full catalogues on brand

Use the brand as the blocking key: the first word of the Abt name, and the first token of the Buy manufacturer, falling back to the first word of the Buy name when the manufacturer is blank.

Status
Accepted in 2020, recorded retrospectively in October 2026
Decided
2020 Semester 2, Project 2 task 1b (coursework/project-2/task1b.py)
Recorded
2026-10-06, while reviving the coursework

Context

Task 1b used the full catalogues: 1,081 Abt and 1,092 Buy products, so 1,180,452 possible pairs, with 1,097 true matches. Instead of linking, it asked for a blocking scheme: assign every record to one or more blocks so that only records sharing a block are compared later. A staff script scored it by pair completeness (PC, the share of true matches that share a block) and reduction ratio (RR, the share of all pairs that no longer need comparing). The two pull against each other: one giant block keeps every match (PC 1) and saves nothing (RR 0).

Decision

Use the brand as the blocking key: the first word of the Abt name, and the first token of the Buy manufacturer, falling back to the first word of the Buy name when the manufacturer is blank.

Options considered

  • No blocking. PC 1 and RR 0: every one of the 1.18 million pairs compared.
  • First word of the name on both sides. Ignores the Buy manufacturer field, which is the more reliable brand when it is present.
  • Brand plus product line (first two words). Much smaller blocks, but it splits most true matches because the two shops word product lines differently.
  • Sorted-neighbourhood or q-gram blocking. More robust to spelling, but more machinery than the task needed, and not tried.
  • Brand (chosen). Every product has one, and the two shops mostly agree on it.

Why

A product's brand is almost always the same in both shops, and it splits the catalogues into about 85 shared blocks of modest size, so most of the comparison work disappears while nearly all matches survive. It also needs no tuning.

What happened

  • As reported: PC 0.9572 and RR 0.9452 (1,050 of 1,097 true matches share a block; 64,723 comparisons instead of 1,180,452). The revival adds a 95% Wilson interval for PC of [0.944, 0.968] and a record-resampling bootstrap interval for RR of [0.941, 0.950] (1,000 resamples, seed 2020).
  • The 47 missed matches are almost all brand aliases and spellings: "tomtom" against "tom", "polk" against "polkaudio", "tripp-lite" against "tripp", a misspelt "sennheisser", and parent companies such as "maytag" against "whirlpool" or "audiovox" against "xm". Blocking on brand cannot recover these, and no later step can either.
  • The 2026 explorer compares other keys on the same scale. Keeping only the first three characters of the brand reaches PC 0.973 [0.961, 0.981] for an RR of 0.945, almost exactly the same cost: with hindsight, a strictly better choice that I did not test in 2020. Coarser prefixes buy little more completeness for a lot more work (one character: PC 0.975, RR 0.867).
  • The staff script reads block files with pandas, which silently turns keys such as "nan" or "null" into missing values. The port reproduces that so the reported numbers match exactly.

What I'd change

  • Normalise brands first: lower-case, strip punctuation, and map known aliases and parent companies with a small, reviewed alias table.
  • Block on more than one key (brand, and separately a normalised model code) and take the union, so a pair only needs to agree on one of them.
  • Plot the PC/RR trade-off for several candidate keys before choosing one, as the 2026 explorer now does, rather than picking the first sensible key.