Skip to content
Data Processing LabUniMelb · 2020 S2

Project 2 · Part 2 · Classification

Reading a country's life expectancy

Twenty World Development Indicators for 2016 (electricity access, fertility, health spending and so on) for 183 countries, each labelled High, Medium or Low life expectancy. Task 2a compares a decision tree with k-nearest neighbours; task 2b engineers 211 features and asks which four describe a country best. Every model is re-trained in your browser by a TypeScript port of the scikit-learn code paths the original used.

decision tree
0.745decision tree
k-NN, k = 3
0.655k-NN, k = 3
feature eng. (2b)
0.729feature eng. (2b)
Task 2a

Decision tree versus k-NN

Missing values (“..” in the CSV) are replaced by each indicator's median, the table is scaled, and 70% of countries train the models. The report found the depth-3 tree (0.745) ahead of 3-NN (0.655) and 7-NN (0.636). Change the split to see how much of that gap is luck.

Training k-NN and decision trees in a background worker…

Data source: World Development Indicators, World Bank (CC BY 4.0); 2016 values for 20 indicators, merged with a life-expectancy class; missing values marked.

Task 2b

Which four features?

Feature engineering multiplies every pair of standardised indicators and adds a k-means cluster label, then keeps the four with the highest ANOVA F-score. PCA keeps four principal components instead, and the baseline simply takes the first four raw columns.

Building 211 features, clustering, selecting and projecting in a background worker…

How much do four components keep?

Explained-variance ratio of the leading principal components of the 211 engineered features (the matrix the original PCA ran on). The four components the model used keep 81.5% of the variance. Under the first four: 95% percentile-bootstrap intervals from 200 resamples of the 183 countries (seed 2020), holding the feature construction fixed. Resampling tends to inflate the largest shares, so read them as stability bands rather than exact intervals.

  • PC 142.6% · cum. 43%95% CI 41.3% to 52.7%
  • PC 231.8% · cum. 74%95% CI 21.5% to 34.4%
  • PC 34.2% · cum. 78%95% CI 3.3% to 7.1%
  • PC 43.0% · cum. 81%95% CI 2.6% to 4.2%
  • PC 52.8% · cum. 84%
  • PC 62.6% · cum. 87%
  • PC 72.0% · cum. 89%
  • PC 81.3% · cum. 90%
  • PC 91.2% · cum. 91%
  • PC 101.0% · cum. 92%

The 211 columns were not re-standardised after the pairwise products were formed, so the 190 products carry about 93% of the total variance and the 20 standardised indicators only about 7%; the largest products pair income and health-spending indicators. The components therefore mostly describe the products. That is the 2020 pipeline as submitted; standardising the 211 columns before PCA is one of the changes listed on the methods page.

View the spectrum as a table
ComponentRatio95% CICumulative
PC 10.4255[0.4126, 0.5267]0.4255
PC 20.3176[0.2155, 0.3441]0.7431
PC 30.0417[0.0328, 0.0712]0.7848
PC 40.0302[0.0260, 0.0419]0.8150
PC 50.0279–0.8429
PC 60.0257–0.8686
PC 70.0201–0.8886
PC 80.0133–0.9020
PC 90.0123–0.9143
PC 100.0100–0.9243
Added in 2026

Is the tree really better?

The 2020 report compared single accuracies from one 70/30 split. This section keeps those numbers and asks how much they would move: Wilson intervals and an exact McNemar test on the submitted split, then the same three models under repeated stratified cross-validation with a paired, corrected comparison. Method and caveats are on the methods page; the model card records the intended use and known failure modes.

Running repeated stratified cross-validation in a background worker…

Back to the start

How the coursework was revived, and what changed.

About this project