A card for the classifiers in Project 2 task 2a (COMP20008, 2020): a decision tree, compared with k-nearest neighbours, that sorts countries into High, Medium or Low life expectancy from World Development Indicators. The 2020 numbers are kept exactly as reported; the intervals and the cross-validation were added in 2026.
Model details
| Field | Detail |
|---|---|
| Model | Decision tree, max_depth=3, random_state=200 (the submitted model); k-NN with k = 3 and k = 7 for comparison |
| Inputs | 20 World Development Indicators for 2016 (for example access to electricity, fertility rate, health spending per person), median-imputed and standardised |
| Output | One of three classes: High, Medium or Low life expectancy at birth |
| Built with | scikit-learn 1.1.3 in 2020; re-run in the browser by a TypeScript port that reproduces every tree node and prediction |
| Author | Sunchuangyu (Rin) Huang, individual coursework, 2020; revived 2026 |
| Code | coursework/project-2/task2a.py (original), web/src/lib/wdi.ts and web/src/lib/ml/ (port), web/src/lib/learning-rigour.ts (2026 evaluation) |
Intended use
- Intended: teaching and portfolio demonstration of a small classification workflow, and of how much one train/test split can mislead.
- Not intended: any decision about a country, a population or a person: health policy, funding, risk scoring, rankings or forecasting. The model describes 2016 cross-sectional data; it does not explain causes and it has not been validated for any operational purpose.
Training data
- World Bank World Development Indicators, 2016 values (CC BY 4.0), joined by
the course to a three-class life-expectancy label (
life.csv). After the join, 183 countries remain: 76 High, 53 Medium and 54 Low. - The cut-offs that define High, Medium and Low were set by the course and are not documented in the files I have.
- Missing values (".." in the CSV) are replaced by each indicator's median. Most indicators are complete; four have 13 to 21 missing countries (adjusted net national income per person, rural and urban drinking water, urban sanitation). Medium-class countries have about twice as many missing values on average (1.0 per country) as High or Low (0.5).
- See the data statement (
docs/data-statement.md) for provenance and licensing.
Evaluation
All intervals are 95%.
The submitted split (seed 200: 128 training and 55 test countries), Wilson intervals:
| Model | Accuracy (as reported) | 95% interval |
|---|---|---|
| Decision tree, depth 3 | 0.745 (41 of 55) | [0.617, 0.842] |
| k-NN, k = 3 | 0.655 (36 of 55) | [0.523, 0.766] |
| k-NN, k = 7 | 0.636 (35 of 55) | [0.504, 0.751] |
On the same 55 countries the tree was right where 3-NN was wrong 11 times, and the reverse 6 times: exact McNemar p = 0.33. The 2020 report's conclusion that the tree was better rested on one split that cannot separate the models.
Repeated stratified cross-validation (10 repeats of 5-fold, seed 2020, imputation and scaling fitted inside each training fold), with Nadeau-Bengio corrected intervals:
| Model | Mean accuracy | 95% interval |
|---|---|---|
| Decision tree, depth 3 | 0.750 | [0.685, 0.816] |
| k-NN, k = 3 | 0.721 | [0.645, 0.798] |
| k-NN, k = 7 | 0.736 | [0.667, 0.805] |
Paired on identical folds, the tree's advantage over 3-NN is +2.9
percentage points, 95% interval [−4.4, +10.2], p = 0.43 (it won 28 of 50
folds, lost 11, tied 11). Against 7-NN it is +1.4 points [−5.6, +8.5],
p = 0.69. These numbers are pinned by a unit test
(web/src/lib/rigour.test.ts), so this card cannot drift from the code
without the test failing.
Per class (submitted split, tree): High 17 of 23 right, Medium 14 of 21, Low 10 of 11. Errors are almost all between neighbouring classes; no High country was predicted Low or the reverse.
How it decides
The fitted tree's first split is the share of deaths from communicable, maternal, prenatal and nutritional conditions (at about 27%). Countries above that share go straight to a leaf; below it, the next split is urban sanitation access (at about 98%). Thresholds are shown in each indicator's own units on the site.
Known failure modes and limitations
- Near-proxy features. The root split is a cause-of-death share, and other inputs (fertility, lifetime risk of maternal death) are closely tied to life expectancy itself. The model mostly re-describes health and demographic correlates of the label rather than learning anything surprising.
- Medium is the hard class. It sits between the other two, has the most errors, and has the most imputed values, so median imputation pulls its countries towards the middle of the feature space.
- Small sample. 183 countries, 55 in the submitted test set. Single-split accuracies move by several points with the seed, which is what the lab's re-split control and the cross-validation show.
- Leakage in the 2020 pipeline. Medians were computed on the whole table before splitting. The effect is small, and the 2026 cross-validation fits them inside each fold, but the reported 2020 numbers include it.
- One year. 2016 values only; nothing here says how the model would do on other years.
Ethical considerations
- Country-level labels invite ranking and stereotyping. The classes describe aggregate outcomes in 2016, not the people who live there.
- Inferring anything about individuals from these country-level patterns would be an ecological fallacy.
- Data quality and missingness differ between national statistical systems, and missing values are not random: imputing medians treats countries with weaker statistical systems as more "average" than they may be.
- Nothing in the model is causal. A split on an indicator says the indicator separates the classes in this table, not that changing it would change life expectancy.
What I'd change
- Report the cross-validated comparison as the headline and the single split only as an example.
- Impute inside the training data from the start, and add a missingness indicator for the four indicators with gaps.
- Drop or flag near-proxy features and see how much accuracy is left.
- Use an ordinal model, because High, Medium and Low are ordered.