Skip to content
Data Processing LabUniMelb · 2020 S2

Data statement

What data, from where, and what is left out

Two public datasets are published in derived form; one copyrighted source is kept out entirely.

Abt-Buy
Leipzig University benchmark
WDI
World Bank, CC BY 4.0
BBC text
Excluded (copyright)

What data this project uses, where it came from, what the site publishes, and what it deliberately leaves out.

Abt-Buy product matching benchmark

  • Source: the Abt-Buy entity-resolution benchmark from the Leipzig University database group (https://dbs.uni-leipzig.de/research/projects/benchmark-datasets-for-entity-resolution), in the subset distributed with COMP20008 in 2020.
  • What it is: product listings from two online electronics shops (Abt.com and Buy.com), with a hand-made list of which listings describe the same product. Records carry an id, a name, a description, and for Buy a manufacturer (and a price, which the site drops).
  • What the site uses: task 1a's 248 Abt and 249 Buy records with their 149 true matches (web/public/data/abt-buy-small.json), and the names and manufacturers of task 1b's 1,081 Abt and 1,092 Buy records with their 1,097 true matches (abt-buy-full.json). Both are produced from the coursework files by scripts/derive_linkage.py.
  • Why it is published: the record-linkage lab and the LLM evaluation run the matcher on the records themselves in the visitor's browser.
  • Licence: no licence is recorded in the coursework files. The site credits the source on every page that uses it; whether redistributing this subset is acceptable was raised as an owner decision when the site was first published and is still open.
  • Known issues: the ground truth is the benchmark's, not mine, and it has the usual imperfections of hand-labelled matches. Descriptions contain marketing text and encoding artefacts (the original code reads them as ISO-8859-1).
  • Sent to an AI provider? Only if a visitor runs the LLM judge on their own key: then the names, the first 500 characters of the descriptions and the Buy manufacturer of each sampled pair go to the provider they chose. No ids or ground-truth labels are sent.

World Development Indicators

  • Source: World Bank, World Development Indicators (https://datacatalog.worldbank.org/search/dataset/0037712/world-development-indicators), licensed CC BY 4.0. The site's copy is modified: 2016 values for 20 indicators, joined to a three-class life-expectancy label supplied with the course (life.csv), with missing values marked.
  • What the site uses: the merged 183-country, 20-indicator table as the original task2a.py builds it (web/public/data/wdi-life.json), produced by scripts/derive_learning.py.
  • Known issues: values are national aggregates of varying quality, and missing values are more common for some groups of countries than others (see the model card). The class cut-offs are the course's and are not documented in the files.
  • Sent to an AI provider? No. No AI feature uses this dataset.

BBC Sport match reports (excluded)

  • What it was: Assignment 1 crawled 147 rugby union match reports, BBC Sport articles mirrored on a course server that no longer exists.
  • What is excluded: the article text and the headlines. They are BBC copyright. The scraped text (a1-articles.csv) is git-ignored and stays on the author's machine; the site never shows headlines or text.
  • What is published instead: a text-free derivative: crawl order, page id, URL, position in the site's link ring, the extracted team and the score under both score rules (coursework/assignment-1/a1-articles-derived.csv and web/public/data/rugby.json). See decision record DR-003.
  • Automated check: web/src/lib/data-exclusion.test.ts fails if any of the 147 headlines appears in the published data files or the documentation.
  • Open item: the original outputs coursework/assignment-1/task1.csv and task2.csv contain the headlines. The repository is private; whether to keep them before making it public is an owner decision recorded in DR-003.

What is not here

  • No personal information about anyone.
  • No data from any employer or workplace system.
  • No AI-generated content in the published data: AI answers exist only in a visitor's own browser, in the AI log and in files they export.