Context
Assignment 1 crawled 147 BBC Sport rugby match reports from a course server,
then extracted the first team mentioned and the score from each article. The
server no longer exists, so reproducing tasks 2 to 5 needs the scraped text,
which I still have locally as a1-articles.csv. That text, and the headlines,
are BBC copyright. The original outputs task1.csv and task2.csv contain the
headlines because that is what the task asked for.
Decision
Publish only a text-free derivative of the articles, and keep the scraped
text out of the repository and off the site. The web app's rugby.json and
coursework/assignment-1/a1-articles-derived.csv hold the crawl order, page
id, URL, position in the site's link ring, the extracted team and the score
under both score rules: everything the analysis needs, and no headline or
article text. a1-articles.csv is git-ignored and stays on my machine;
scripts/derive_rugby.py regenerates the derivative from it.
Options considered
- Publish the article text. The simplest way to make the lab fully re-runnable, but it is not mine to publish.
- Publish headlines only. Less text, still copyrighted, and not needed by any analysis on the site.
- Write synthetic articles. Would let visitors re-run extraction end to end, but would quietly change what the 2020 numbers were computed on.
- Drop the rugby lab. Avoids the question by losing a third of the coursework.
- Derived fields only (chosen).
Why
Tasks 2 to 5 only depend on the team and score extracted from each page, and task 1 only on the link structure. Publishing those keeps every reported number reproducible and testable (the TypeScript port is checked against the derivative), while respecting the copyright of the source. The "paste a match report" panel lets visitors try the original extraction logic on text they supply themselves.
What happened
- The rugby lab runs entirely from a 15 KB
rugby.json. When the site first went live I checked that none of the 147 headlines appear in the deployed HTML, the data files or the JavaScript bundles; since October 2026 a unit test (web/src/lib/data-exclusion.test.ts) repeats that check for the data files and the documentation on every CI run. - The limitation is real: nobody else can re-run
derive_rugby.py, because it needs the text. They can only check that the derivative reproduces the reported tables. coursework/assignment-1/task1.csvandtask2.csvstill contain the headlines, because thecoursework/folder preserves the submission unchanged. While the repository is private that is acceptable; it is an open question for the owner before the repository is made public.
What I'd change
- Decide the
task1.csv/task2.csvquestion before going public: either keep the repository private, or replace the headline column with a stable hash and say so incoursework/README.md, accepting that those two files are then no longer byte-for-byte originals. - Extend the automated check to the built JavaScript bundles as well as the data files.