Guided tour · recorded from this site
Watch the labs at work
Three short walkthroughs follow the main workflows from start to finish, each step shown on screen and as captions. A Playwright script recorded them from a production build of this site and checks every step as it goes (the extracted score under both rules, the reported precision and recall, the blocking numbers), so the same inputs reproduce the same numbers.
Mine a match report
A short synthetic match report pasted into the Assignment 1 extractor: string search finds the team, the original regular expression finds the scores, and the per-team charts from all 147 crawled pages follow.
Steps (transcript)
- 1Open the Rugby Report Miner: the 2020 extractor, ported line for line
- 2Paste a short match report, written for this demo
- 3Team names are marked; Scotland is named first, so the report is about Scotland
- 4Score tokens are marked: 27-20 has the largest total, so it is kept
- 52014-15 is longer than five characters once cleaned, so the original ignores it
- 6The final script's rule keeps the largest single number instead: 30-3
- 7Per-team charts from all 147 crawled pages: articles per team
- 8Average winning margin, with n and a 95% interval under each bar
- 9Coverage against margin, one point per team
Link records
The four-pass Abt-Buy matcher with its pass 3 threshold moved by hand, the matched pairs checked against the ground truth, then the blocking trade-off on the full catalogues.
Steps (transcript)
- 1Open Record Linkage: the same products, listed by two shops in different words
- 2As reported: precision 0.864 and recall 0.852, each with a Wilson 95% interval
- 3Lower the pass 3 threshold to 0.30: recall rises to 0.899, precision falls to 0.779
- 4Raise it to the script's 0.50: precision 0.887, recall 0.792
- 5Every point on the sweep is a full re-run of the four passes, with Wilson bands
- 6Matched pairs, checked against the 149 true matches: show the wrong ones
- 7Blocking the full catalogues on brand keeps 95.7% of true matches
- 8Other keys move along the trade-off: completeness against comparisons left
- 9No blocking keeps every match but compares all 1.18 million pairs
Classical matcher vs LLM
The evaluation harness: the 2020 matcher and an LLM judge on the same 40 sampled pairs, scored with intervals and a paired test. Here the judge is a labelled mock; no real key or model is used.
Mocked AI response for illustration. Steps 5 to 10 use a placeholder key and the model id mock-for-illustration. Requests to the provider are intercepted in the browser and answered by a mock that calls a pair a match when both names share a model number; no model was called, so the LLM column shows the mock, not a real model's results.
Steps (transcript)
- 1Open the LLM evaluation: the 2020 matcher and an LLM judge, scored on the same pairs
- 240 pairs, seed 2020, half of them matches: the matcher's scores with 95% intervals
- 3AI settings: bring your own key, kept in this browser and sent only to the provider
- 4For this demo: OpenAI, a placeholder key and the model id “mock-for-illustration”
- 5Run the judge on the 40 pairs: every answer is schema-checked and loggedMocked AI response for illustration
- 6Side by side: precision, recall and F1 with Wilson and bootstrap 95% intervalsMocked AI response for illustration
- 7Two confusion matrices on the same pairsMocked AI response for illustration
- 8Paired comparison: McNemar's exact test and the difference with its intervalMocked AI response for illustration
- 9Read the disagreements: each answer is labelled AI-generated; accept, flip or reject itMocked AI response for illustration
- 10The AI log: every call without the key, with the review decision and JSON or CSV exportMocked AI response for illustration
- 11Forget key: the placeholder is removed from this browser
Every key feature at a glance
Captured by the same script in light mode at 1440 × 900 (the landing page in dark mode too) and on a 390 px phone. Select one to enlarge it; the arrow keys step through the set.
Desktop · 1440 × 900
Mobile · 390 × 844
How these were made
pnpm showcase runs web/e2e/showcase.spec.ts on the system Chrome. It plays each journey at a human pace with an on-screen caption and cursor, asserts what it shows, and records it at 1280 × 800; ffmpeg then encodes the H.264 videos on this page and the GIFs in the README. The step lists and captions here are the same text as the on-screen steps. The match report in the first walkthrough was written for the demo; the 2020 article text is not published.
No real API key appears anywhere in these recordings. Where the AI feature is shown, the key is a placeholder, every request to the provider is intercepted in the browser, and the reply comes from a labelled mock with a fixed rule, so its numbers say nothing about any real model. To see a real run, add your own key on the LLM evaluation page.