Skip to content
Data Processing LabUniMelb · 2020 S2

Decision record · DR-003

Publish only derived rugby data, never the article text

Publish only a text-free derivative of the articles, and keep the scraped text out of the repository and off the site.

Status
Accepted
Decided
2026, while reviving Assignment 1 for the web
Recorded
2026-10-06

Context

Assignment 1 crawled 147 BBC Sport rugby match reports from a course server, then extracted the first team mentioned and the score from each article. The server no longer exists, so reproducing tasks 2 to 5 needs the scraped text, which I still have locally as a1-articles.csv. That text, and the headlines, are BBC copyright. The original outputs task1.csv and task2.csv contain the headlines because that is what the task asked for.

Decision

Publish only a text-free derivative of the articles, and keep the scraped text out of the repository and off the site. The web app's rugby.json and coursework/assignment-1/a1-articles-derived.csv hold the crawl order, page id, URL, position in the site's link ring, the extracted team and the score under both score rules: everything the analysis needs, and no headline or article text. a1-articles.csv is git-ignored and stays on my machine; scripts/derive_rugby.py regenerates the derivative from it.

Options considered

  • Publish the article text. The simplest way to make the lab fully re-runnable, but it is not mine to publish.
  • Publish headlines only. Less text, still copyrighted, and not needed by any analysis on the site.
  • Write synthetic articles. Would let visitors re-run extraction end to end, but would quietly change what the 2020 numbers were computed on.
  • Drop the rugby lab. Avoids the question by losing a third of the coursework.
  • Derived fields only (chosen).

Why

Tasks 2 to 5 only depend on the team and score extracted from each page, and task 1 only on the link structure. Publishing those keeps every reported number reproducible and testable (the TypeScript port is checked against the derivative), while respecting the copyright of the source. The "paste a match report" panel lets visitors try the original extraction logic on text they supply themselves.

What happened

  • The rugby lab runs entirely from a 15 KB rugby.json. When the site first went live I checked that none of the 147 headlines appear in the deployed HTML, the data files or the JavaScript bundles; since October 2026 a unit test (web/src/lib/data-exclusion.test.ts) repeats that check for the data files and the documentation on every CI run.
  • The limitation is real: nobody else can re-run derive_rugby.py, because it needs the text. They can only check that the derivative reproduces the reported tables.
  • coursework/assignment-1/task1.csv and task2.csv still contain the headlines, because the coursework/ folder preserves the submission unchanged. While the repository is private that is acceptable; it is an open question for the owner before the repository is made public.

What I'd change

  • Decide the task1.csv / task2.csv question before going public: either keep the repository private, or replace the headline column with a stable hash and say so in coursework/README.md, accepting that those two files are then no longer byte-for-byte originals.
  • Extend the automated check to the built JavaScript bundles as well as the data files.