Methods · decisions · cards
How it was built, and how far to trust it
Where the data came from, what each lab does, how every number gets its interval, what is assumed, what is weak, and the decisions behind it, written down in one place. The 2020 results stay exactly as reported; everything added in 2026 sits around them.
Three sources, one left out
Every published file is derived from the original coursework by a script that first checks the original code still reproduces the reported numbers. The data statement has the full provenance and licences.
Abt-Buy benchmark
Product listings from two shops with hand-labelled matches (Leipzig University database group), in the course's subset: 248 × 249 records for linkage, 1,081 × 1,092 for blocking.
web/public/data/abt-buy-small.json
World Development Indicators
World Bank, CC BY 4.0: 20 indicators for 2016 joined to a course-supplied High / Medium / Low life-expectancy label, 183 countries, missing values marked.
web/public/data/wdi-life.json
BBC match reports (excluded)
147 copyrighted articles crawled in 2020. Only a text-free derivative is published (crawl order, team, score); a unit test fails if any headline reaches the site's data.
coursework/assignment-1/a1-articles-derived.csv
Faithful ports, then a layer of rigour
The 2020 Python is ported to TypeScript and tested against outputs produced by running the untouched originals again with pinned 2020-era libraries: several hundred parity checks on crawl order, extracted scores, matched pairs, block files, random splits, every tree node and every k-NN prediction. The 2026 additions never change those numbers; they add intervals, paired tests and an evaluation harness around them.
- Rugby report minerOpen
- Breadth-first crawl of a reconstructed link ring; the first team mentioned and the largest score found by the original string search and regular expression; per-team article counts and mean winning margins, now with bootstrap intervals.
- Record linkageOpen
- The four-pass matcher (DR-001) and brand blocking (DR-002), scored by the staff measures; Wilson intervals on precision, recall and pair completeness, bootstrap intervals on F1 and the reduction ratio, and Wilson bands across the pass-3 threshold sweep.
- ClassificationOpen
- Median imputation, scaling, a depth-3 decision tree and k-NN on a 70/30 split, then interaction features, k-means, SelectKBest and PCA. Added: Wilson intervals on every accuracy, repeated stratified CV with a corrected paired comparison, and the PCA variance spectrum.
- LLM evaluationOpen
- A seeded, stratified, record-disjoint sample of candidate pairs judged by the 2020 matcher (run on the whole catalogue) and, on the visitor's own key, an LLM (shown one pair at a time); scored side by side with paired tests, and split by whether the matcher's verdict leaned on catalogue-wide context (DR-004).
Every number gets an interval
All intervals are 95%. The helpers live in web/src/lib/stats and are unit-tested against scipy and statsmodels values generated by scripts/stats_reference.py. Seeds are fixed and printed next to each result: 2020 for everything added in 2026, and the originals' own 200 (task 2a split) and 12517 (task 2b splits).
| Quantity | Interval or test | Where |
|---|---|---|
| Proportions: precision, recall, accuracy, specificity, pair completeness | Wilson score interval | Linkage, blocking, classification, LLM evaluation |
| Linkage F1 on the full task 1a run | Percentile bootstrap resampling Abt records (2,000, seed 2020) | Linkage |
| Reduction ratio | Percentile bootstrap resampling Abt and Buy records (1,000, seed 2020) | Blocking |
| A team's average winning margin | Percentile bootstrap over its articles (2,000, seed 2020) from 8 articles; below that a Student t interval truncated at 0, because a bootstrap of two to seven values is far too narrow | Rugby |
| Cross-validated accuracy and paired differences | Nadeau-Bengio corrected resampled t (10 × 5-fold, seed 2020) | Classification |
| Two classifiers on one test set | Exact McNemar test with the conditional odds ratio | Classification, LLM evaluation |
| F1 and LLM-minus-matcher differences on a sample | Stratified, paired percentile bootstrap (2,000, seed 2020) | LLM evaluation |
| Precision at the candidate pool's match rate | Conservative bounds from 97.5% Wilson intervals on recall and specificity | LLM evaluation |
| PCA explained-variance shares | Percentile bootstrap resampling countries (200, seed 2020), computed at build time | Feature engineering |
Paired, not side by side
When two methods are compared they are scored on the same units (the same 55 countries, the same 50 folds, the same sampled pairs), and the comparison is made on the differences: McNemar's test on the units where only one method was right, and an interval for the mean difference, which is reported as the effect size in the metric's own units rather than as a p-value alone.
What the intervals changed
The 2020 report preferred the decision tree (0.745) to 3-NN (0.655). On the same split the exact McNemar p is 0.33, and across 10 × 5-fold CV the tree's lead is 2.9 points with a 95% interval of −4.4 to +10.2: the data do not separate the two. The reported numbers stand; the conclusion drawn from them does not.
What has to be true, and what is weak
Assumptions
- The Abt-Buy ground truth is correct. Every precision and recall on the site, 2020 or 2026, is measured against it.
- The 147 crawled pages are the whole rugby site. The crawl is replayed on a reconstructed link ring that reproduces the saved visit order exactly; the original server is gone.
- The 183 countries are treated as a sample when intervals are computed, although they are close to the full population of countries. The intervals describe how much the numbers depend on which countries happen to be in the table, not sampling error in the usual sense.
- Pairs, records and countries are treated as independent units in the bootstrap and Wilson intervals. Where that clearly fails (one Abt record's predicted and true pairs in the full linkage run), the resampling unit is the record. An LLM-evaluation sample is drawn so that no Abt or Buy listing appears in two of its pairs, so its 20 to 120 pairs are resampled as pairs. The matcher's verdicts still depend on records outside the sample, through its one-to-one rule.
- Seeds are fixed so that every number on the site can be reproduced; changing a seed is offered wherever the result depends on one.
Limitations
- Small samples throughout: 55 test countries in the submitted split, 149 true matches in task 1a, four articles behind New Zealand's average margin. Many intervals are wide, and the site shows them rather than hiding them.
- The cross-validation folds are seeded but not bit-identical to scikit-learn's RepeatedStratifiedKFold; the per-split ports (train_test_split, the tree builder, k-NN) are bit-identical.
- Re-weighting the LLM evaluation's precision to the pool's 4.9% match rate needs many more non-matches than a 40-pair sample holds, so that interval is wide by design.
- The LLM evaluation is not like for like: the matcher's verdict on a pair uses the whole catalogue (a listing it has already linked is skipped, and passes 3 and 4 keep only the best candidate), while the LLM sees one pair. In the default sample, 15 of the 20 non-matches involve a listing the matcher had linked elsewhere. The page splits both methods' scores by that context instead of hiding it.
- No LLM results are published: there was no budget to run the judge, and numbers a visitor cannot reproduce would undercut the point of the harness.
- The rugby lab cannot be re-run end to end by anyone else, because the source articles are copyrighted and kept out of the repository (DR-003).
What I'd change
- Drop the matcher's fourth pass, which found 8 pairs and got 1 right, and replace the greedy pass order with a one-to-one assignment. DR-001
- Normalise brands and block on brand and model code together; even a three-character brand prefix keeps 17 more true matches for almost no extra comparisons. DR-002
- Report cross-validated comparisons as the headline, impute inside training folds, and drop near-proxy features from the life-expectancy model.
- Standardise the 211 engineered features before PCA; as submitted, the 190 pairwise products carry about 93% of the variance.
- Evaluate the LLM judge with far more non-matches than matches, on a fresh sample after any prompt change. DR-004
- Decide whether the original task1.csv and task2.csv, which hold BBC headlines, can stay before the repository is made public. DR-003
AI use statement
Also in docs/ai-use-statement.md. Every AI call made from your browser is listed in the AI log.
This statement covers the optional AI feature on this site. It is informed by the Australian Government's policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It is a portfolio demonstration and makes no claim of compliance or certification against any of them.
What the AI does
One feature uses a large language model: the LLM judge on the LLM
evaluation page (/llm-eval). For each sampled pair of product listings it
answers whether the two listings describe the same product, with a confidence
(low, medium or high) and a one or two sentence rationale. Its answers are
scored against the Abt-Buy ground truth next to the 2020 matcher.
What it never does
- It never runs unless a visitor adds their own API key and presses Run.
- It never changes the 2020 results, the matcher, or any number reported from the original coursework.
- It never decides anything for anyone: it is an object of evaluation, and its answers are shown as such.
- It is not used anywhere else on the site, and nothing on the site was written by it at run time.
Data sent to the provider
Only when a visitor runs the judge, and only to the provider they chose
(Anthropic's api.anthropic.com or OpenAI's api.openai.com), directly from
their browser:
- fixed instructions describing the task;
- for each sampled pair: both product names, the first 500 characters of each description, and the Buy manufacturer.
No record ids, ground-truth labels, personal information or anything about the visitor is sent. The visitor's API key goes to the provider as the authentication header and nowhere else: the site has no server, and the key is never written to the AI log, an export or a URL. By default the key is kept in the tab's session storage and disappears when the tab closes; visitors can opt in to keeping it in local storage (and opt out again, which moves it back) and can forget their keys for both providers at any time. The provider's own terms and data practices apply to what it receives.
Models
Claude Haiku 4.5 (claude-haiku-4-5, the default, run at temperature 0) or
Claude Sonnet 5.5 (claude-sonnet-5-5, run at low effort), through
Anthropic's Messages API; or any OpenAI Chat Completions model the visitor
names. Answers must match a JSON schema and are validated before use; anything
that does not validate is recorded as a failed call. No fallback model is
used, so the model recorded is the model that answered.
Transparency and audit
- Every AI answer is labelled AI-generated where it appears.
- Every call that leaves the visitor's browser is appended to an audit log there (IndexedDB): id, time, feature, provider, model, the exact prompt and parameters sent, the answer (or the error), latency, the token usage the provider reported, and the human decision. A call the visitor stops while it is in flight is logged too, as an error, because the provider may still have processed and billed it. If the log cannot be written, the run stops and any unlogged answer is discarded.
- The log is viewable at
/ai-log, filterable, and exportable as JSON or CSV.
Human in the loop
Each AI answer can be accepted, flipped (edited) or rejected by the person reviewing it, and that decision is written to the audit log beside the call. The evaluation scores always use the model's own answers, so a reviewer's decisions never silently improve the model's numbers.
Limits to keep in mind
- Results depend on the prompt, the model version and the sample; forty pairs give wide intervals, which the page shows rather than hides.
- The comparison is not like for like. The judge sees one pair at a time; the 2020 matcher's verdicts come from a run over the whole catalogue, in which a listing already linked is skipped. The page reports both methods' scores split by that context.
- Models can be confidently wrong, and their rationales can sound more certain than the evidence supports. Read the disagreements, not just the totals.
- Costs and rate limits are the visitor's own; the page estimates the cost before a run and reports actual token use after it.
The decisions, written down
Each record states the decision first, the options, why, what happened (weak numbers included) and what I would change. Past records are never rewritten: a new record supersedes an old one.
- DR-001A four-pass matcher for Abt-Buy record linkageLink records in four passes of decreasing strictness, where each pass only considers records the earlier passes left unmatched.Accepted in 2020, recorded retrospectively in October 2026
- DR-002Blocking the full catalogues on brandUse the brand as the blocking key: the first word of the Abt name, and the first token of the Buy manufacturer, falling back to the first word of the Buy name when the manufacturer is blank.Accepted in 2020, recorded retrospectively in October 2026
- DR-003Publish only derived rugby data, never the article textPublish only a text-free derivative of the articles, and keep the scraped text out of the repository and off the site.Accepted
- DR-004How to evaluate an LLM judge against the 2020 matcherEvaluate the matcher and an LLM judge as a paired comparison on a seeded, stratified sample of candidate pairs, with every LLM call made on the visitor's own key and written to an audit log.Accepted