Skip to content
Data Processing LabUniMelb · 2020 S2

COMP20008 · University of Melbourne · 2020 S2

Crawl.
Link.
Learn.

Three pieces of data-processing coursework from Elements of Data Processing: a web crawler that mines rugby match reports, a record linker that finds the same product in two shops, and classifiers that read life expectancy from World Bank indicators. The original Python is ported to TypeScript, checked against its 2020 outputs, and left for you to poke at.

Or watch the guided tour: three workflows, start to finish

Final score · 2020

as reported, reproduced in the labs

Pages crawledkept with team + score
147·64
Linkage P / Rtask 1a, 248 × 249
.864·.852
Blocking PC / RRtask 1b, 1,081 × 1,092
.957·.945
Tree vs 3-NNtask 2a accuracy
.745·.655

The labs

What the coursework asked, and what you can do with it

Added in 2026

Rigour around the results, and an LLM put to the test

The original results stay exactly as submitted. Around them: intervals on every number, paired comparisons instead of eyeballed gaps, written-down decisions, and an evaluation that checks whether a language model would have done the linkage better. The AI part is optional and runs only on your own key.

Revival notes

Faithful, not improved

The aim was to keep every algorithm exactly as submitted, quirks included. Each TypeScript module is tested against numbers produced by running the original scripts again with uv: crawl order, extracted scores, matched pairs, block files, random splits, every tree node and every k-NN prediction. Re-running them also turned up three things the 2020 reports did not say.

Bit-identical where it counts: numpy's Mersenne Twister, scikit-learn's tree builder, difflib and Levenshtein ratios are all re-implemented.

  1. Two score rules

    The notebook that produced the submitted CSVs keeps the score with the largest points total; the final script keeps the one with the largest single number. They disagree on two of 147 reports. Both are ported.

  2. 0.4 in the report, 0.5 in the code

    The linkage report quotes precision 0.8639 and recall 0.8523. Those come from a 0.4 threshold with fuzzywuzzy's pure-Python backend; the code as submitted says 0.5, and its saved output needs the C extension. The lab offers both.

  3. An unseeded k-means

    Task 2b clustered without a random seed, and cluster numbering leaks into the features. The partition is stable, so numbering clusters by first appearance reproduces 0.729, 0.713 and 0.721 to the digit.

About this project

Team sheet

Subject
COMP20008 Elements of Data Processing
University
The University of Melbourne
Term
2020, Semester 2
Work
Assignment 1 and Project 2, both individual
Author
Sunchuangyu (Rin) Huang
Revived
2026, as an interactive web lab
Data
Abt-Buy product matching benchmark, Leipzig University database groupWorld Development Indicators, World Bank, CC BY 4.0 (modified)

Original stack

2020, Python 3 scripts and notebooks

  • requests + BeautifulSoup crawler
  • re, pandas, matplotlib
  • textdistance + fuzzywuzzy
  • scikit-learn: k-NN, decision tree, k-means, PCA, SelectKBest

Revived stack

2026, static site, all compute in the browser

  • Next.js 16, React 19, TypeScript
  • Tailwind CSS 4, shadcn/ui on Base UI
  • Hand-written SVG charts, Web Workers
  • Vitest parity tests against uv-run Python

Source, original submission and reports

The 2020 code is preserved unchanged under coursework/ in the project repository, which is currently private.

Academic integrity: this is my own 2020 submission, shared for reference and as a portfolio piece. The assignment specifications and the course's scraped article text are not reproduced here, and the task descriptions are paraphrased. If you are taking COMP20008, please do your own work.