COMP20008 · University of Melbourne · 2020 S2
Crawl.
Link.
Learn.
Three pieces of data-processing coursework from Elements of Data Processing: a web crawler that mines rugby match reports, a record linker that finds the same product in two shops, and classifiers that read life expectancy from World Bank indicators. The original Python is ported to TypeScript, checked against its 2020 outputs, and left for you to poke at.
Or watch the guided tour: three workflows, start to finishFinal score · 2020
as reported, reproduced in the labs
- Pages crawledkept with team + score
- 147·64
- Linkage P / Rtask 1a, 248 × 249
- .864·.852
- Blocking PC / RRtask 1b, 1,081 × 1,092
- .957·.945
- Tree vs 3-NNtask 2a accuracy
- .745·.655
The labs
What the coursework asked, and what you can do with it
Rugby report miner
Crawl a site of 147 match reports breadth-first, find the team and score in each, and summarise every team's coverage and winning margin.
- Replay the crawl queue step by step
- Paste a report and watch the regex decide
- Flip between the two score rules
Record linkage
Match the same products across the Abt and Buy catalogues, then block the full catalogues so far fewer pairs need comparing.
- Tune the matcher's thresholds
- See every right, wrong and missed pair
- Compare blocking keys on one trade-off chart
Classification lab
Predict life expectancy from World Bank indicators with a decision tree and k-NN, then test engineered features against PCA.
- Slide k from 1 to 40
- Grow the tree and read its splits
- Re-split the data with any seed
Added in 2026
Rigour around the results, and an LLM put to the test
The original results stay exactly as submitted. Around them: intervals on every number, paired comparisons instead of eyeballed gaps, written-down decisions, and an evaluation that checks whether a language model would have done the linkage better. The AI part is optional and runs only on your own key.
LLM judge vs the 2020 matcher
An evaluation harness: the same sampled Abt-Buy pairs judged by the four-pass matcher and, on your own API key, a language model. Precision, recall and F1 with intervals, McNemar's test, cost and latency, a disagreement explorer and an audit log.
Run the evaluationHow sure are the 2020 numbers?
Every reported number keeps its value and gains an interval. Repeated cross-validation with a paired, corrected test shows the decision tree's 9-point lead over 3-NN was mostly one lucky split.
See the uncertaintyMethods and decisions
Data provenance, evaluation design, assumptions and limits, four decision records with their weak numbers left in, a model card, a data statement and an AI use statement.
Read the methodsRevival notes
Faithful, not improved
The aim was to keep every algorithm exactly as submitted, quirks included. Each TypeScript module is tested against numbers produced by running the original scripts again with uv: crawl order, extracted scores, matched pairs, block files, random splits, every tree node and every k-NN prediction. Re-running them also turned up three things the 2020 reports did not say.
Bit-identical where it counts: numpy's Mersenne Twister, scikit-learn's tree builder, difflib and Levenshtein ratios are all re-implemented.
Two score rules
The notebook that produced the submitted CSVs keeps the score with the largest points total; the final script keeps the one with the largest single number. They disagree on two of 147 reports. Both are ported.
0.4 in the report, 0.5 in the code
The linkage report quotes precision 0.8639 and recall 0.8523. Those come from a 0.4 threshold with fuzzywuzzy's pure-Python backend; the code as submitted says 0.5, and its saved output needs the C extension. The lab offers both.
An unseeded k-means
Task 2b clustered without a random seed, and cluster numbering leaks into the features. The partition is stable, so numbering clusters by first appearance reproduces 0.729, 0.713 and 0.721 to the digit.
About this project
Team sheet
- Subject
- COMP20008 Elements of Data Processing
- University
- The University of Melbourne
- Term
- 2020, Semester 2
- Work
- Assignment 1 and Project 2, both individual
- Author
- Sunchuangyu (Rin) Huang
- Revived
- 2026, as an interactive web lab
- Data
- Abt-Buy product matching benchmark, Leipzig University database groupWorld Development Indicators, World Bank, CC BY 4.0 (modified)
Original stack
2020, Python 3 scripts and notebooks
- requests + BeautifulSoup crawler
- re, pandas, matplotlib
- textdistance + fuzzywuzzy
- scikit-learn: k-NN, decision tree, k-means, PCA, SelectKBest
Revived stack
2026, static site, all compute in the browser
- Next.js 16, React 19, TypeScript
- Tailwind CSS 4, shadcn/ui on Base UI
- Hand-written SVG charts, Web Workers
- Vitest parity tests against uv-run Python
Source, original submission and reports
The 2020 code is preserved unchanged under coursework/ in the project repository, which is currently private.
Academic integrity: this is my own 2020 submission, shared for reference and as a portfolio piece. The assignment specifications and the course's scraped article text are not reproduced here, and the task descriptions are paraphrased. If you are taking COMP20008, please do your own work.