What data this project uses, where it came from, what the site publishes, and what it deliberately leaves out.
Abt-Buy product matching benchmark
- Source: the Abt-Buy entity-resolution benchmark from the Leipzig University database group (https://dbs.uni-leipzig.de/research/projects/benchmark-datasets-for-entity-resolution), in the subset distributed with COMP20008 in 2020.
- What it is: product listings from two online electronics shops (Abt.com and Buy.com), with a hand-made list of which listings describe the same product. Records carry an id, a name, a description, and for Buy a manufacturer (and a price, which the site drops).
- What the site uses: task 1a's 248 Abt and 249 Buy records with their 149
true matches (
web/public/data/abt-buy-small.json), and the names and manufacturers of task 1b's 1,081 Abt and 1,092 Buy records with their 1,097 true matches (abt-buy-full.json). Both are produced from the coursework files byscripts/derive_linkage.py. - Why it is published: the record-linkage lab and the LLM evaluation run the matcher on the records themselves in the visitor's browser.
- Licence: no licence is recorded in the coursework files. The site credits the source on every page that uses it; whether redistributing this subset is acceptable was raised as an owner decision when the site was first published and is still open.
- Known issues: the ground truth is the benchmark's, not mine, and it has the usual imperfections of hand-labelled matches. Descriptions contain marketing text and encoding artefacts (the original code reads them as ISO-8859-1).
- Sent to an AI provider? Only if a visitor runs the LLM judge on their own key: then the names, the first 500 characters of the descriptions and the Buy manufacturer of each sampled pair go to the provider they chose. No ids or ground-truth labels are sent.
World Development Indicators
- Source: World Bank, World Development Indicators
(https://datacatalog.worldbank.org/search/dataset/0037712/world-development-indicators),
licensed CC BY 4.0. The site's copy is modified: 2016 values for 20
indicators, joined to a three-class life-expectancy label supplied with the
course (
life.csv), with missing values marked. - What the site uses: the merged 183-country, 20-indicator table as the
original
task2a.pybuilds it (web/public/data/wdi-life.json), produced byscripts/derive_learning.py. - Known issues: values are national aggregates of varying quality, and missing values are more common for some groups of countries than others (see the model card). The class cut-offs are the course's and are not documented in the files.
- Sent to an AI provider? No. No AI feature uses this dataset.
BBC Sport match reports (excluded)
- What it was: Assignment 1 crawled 147 rugby union match reports, BBC Sport articles mirrored on a course server that no longer exists.
- What is excluded: the article text and the headlines. They are BBC
copyright. The scraped text (
a1-articles.csv) is git-ignored and stays on the author's machine; the site never shows headlines or text. - What is published instead: a text-free derivative: crawl order, page id,
URL, position in the site's link ring, the extracted team and the score under
both score rules (
coursework/assignment-1/a1-articles-derived.csvandweb/public/data/rugby.json). See decision record DR-003. - Automated check:
web/src/lib/data-exclusion.test.tsfails if any of the 147 headlines appears in the published data files or the documentation. - Open item: the original outputs
coursework/assignment-1/task1.csvandtask2.csvcontain the headlines. The repository is private; whether to keep them before making it public is an owner decision recorded in DR-003.
What is not here
- No personal information about anyone.
- No data from any employer or workplace system.
- No AI-generated content in the published data: AI answers exist only in a visitor's own browser, in the AI log and in files they export.