How good is a model, really? The confusion matrix and everything built on it — accuracy, precision, recall, F1, ROC and AUC — plus the error measures for regression.
Module 5 · Week 5 · Lecture notes by Dr. Abdulkarim Albanna
Core Evaluation ~65 minPrerequisites: Modules 1, 3 and 4. Every number in this module is computed on test data — data the model did not train on (holdout or cross-validation, Modules 1–2).
After training we need to answer three questions: How accurate is the model on new data? Which of several models is best? Is it good enough to use? The answers come from evaluation metrics computed on a validation or test set of labelled examples — never on the training set, which would reward overfitting.
| Problem type | Metrics in this module |
|---|---|
| Classification | Confusion matrix, accuracy, error rate, precision, recall (sensitivity), specificity, F1, Fβ, ROC curve, AUC |
| Regression | MAE, MSE, RMSE, R², adjusted R² |
To estimate them reliably use the holdout method or k-fold cross-validation, and compare classifiers with ROC curves (Section 9).
A confusion matrix compares the predicted classes with the actual ones. Rows are the actual class, columns the predicted class; entry CMi,j counts the examples of class i that the classifier labelled as class j. A good classifier puts almost everything on the diagonal.
A test classifies people as pregnant (positive) or not pregnant (negative):
Totals: P = TP + FN actual positives, N = FP + TN actual negatives; P' = TP + FP predicted positives, N' = FN + TN predicted negatives.
Accuracy (recognition rate) is the fraction of all examples classified correctly; the error rate is the rest:
Accuracy = (6 + 5) / 14 = 11/14 = 78.6%; error rate = 3/14 = 21.4%.
Accuracy is only useful when the classes are balanced. When one class is rare — fraud, cancer, HIV-positive — a model can score very high accuracy while missing almost every case that matters.
| Actual \ Predicted | cancer = yes | cancer = no | Total | Rate |
|---|---|---|---|---|
| cancer = yes | 90 | 210 | 300 | 30.0% (sensitivity) |
| cancer = no | 140 | 9,560 | 9,700 | 98.6% (specificity) |
| Total | 230 | 9,770 | 10,000 | 96.5% (accuracy) |
Cancer example from Han, Kamber & Pei, Data Mining: Concepts and Techniques.
A 96.5% accurate cancer test that finds only 90 of the 300 patients (30%) is a poor test. Predicting “no cancer” for everyone would score 97% accuracy and find none. We need measures that look at each class separately.
Precision = 6 / (6 + 2) = 6/8 = 0.75 — 75% of the points called positive are positive. Recall = 6 / (6 + 1) = 6/7 = 0.857 — the model found 86% of the positives.
Cancer test above: precision = 90/230 = 39.1%, recall = 90/300 = 30.0% — both poor, despite 96.5% accuracy.
Which one matters more depends on the cost of each mistake. A spam filter needs high precision (a real e-mail sent to spam is costly). A cancer screening test needs high recall (a missed patient is costly). Moving the decision threshold trades one for the other (Module 4, Section 7).
Which is better: model 1 with precision 50% and recall 80%, or model 2 with precision 70% and recall 30%? We want one number. The plain average is misleading: a model with precision 100% and recall 0% (it found nothing) averages a respectable 50%.
The harmonic mean fixes this: it stays close to the smaller of the two numbers, so if either precision or recall is small, F1 is small too.
Model 1: F1 = 2(0.5)(0.8) / (0.5 + 0.8) = 0.8 / 1.3 = 0.615. Model 2: F1 = 2(0.7)(0.3) / (0.7 + 0.3) = 0.42 / 1.0 = 0.42. Model 1 is better balanced. Precision 100%, recall 0%: F1 = 0, not 50%. The 14 points: F1 = 2(0.75)(0.857) / 1.607 = 0.80.
Fβ lets you weight one side more: β = 2 counts recall twice as much as precision (screening); β = 0.5 favours precision (spam). β = 1 gives F1.
Dividing each cell of the matrix by its row total gives four rates. Each row's two rates add up to 1.
| Rate | Also called | Formula | Relation |
|---|---|---|---|
| True Positive Rate (TPR) | Sensitivity, recall | TP / P = TP / (TP + FN) | 1 − FNR |
| False Negative Rate (FNR) | Miss rate | FN / P = FN / (TP + FN) | 1 − TPR |
| True Negative Rate (TNR) | Specificity | TN / N = TN / (TN + FP) | 1 − FPR |
| False Positive Rate (FPR) | False-alarm rate, fall-out | FP / N = FP / (FP + TN) | 1 − TNR |
| Measure | Formula | 14 points |
|---|---|---|
| Accuracy | (TP + TN) / (P + N) | 11/14 = 0.786 |
| Error rate | (FP + FN) / (P + N) | 3/14 = 0.214 |
| Sensitivity / recall / TPR | TP / P | 6/7 = 0.857 |
| Specificity / TNR | TN / N | 5/7 = 0.714 |
| Precision | TP / (TP + FP) | 6/8 = 0.750 |
| FPR | FP / N | 2/7 = 0.286 |
| F1 | 2PR / (P + R) | 0.800 |
A model is tested on 200 images: 170 contain cats and 30 do not. It says 160 images contain cats and 40 do not; 140 of its “cat” answers are right.
| Actual \ Predicted | cat | not cat | Total |
|---|---|---|---|
| cat | TP = 140 | FN = 30 | P = 170 |
| not cat | FP = 20 | TN = 10 | N = 30 |
| Total | P' = 160 | N' = 40 | 200 |
Precision = 140/160 = 0.875; recall = 140/170 = 0.824; accuracy = 150/200 = 0.75; error rate = 50/200 = 0.25; specificity = 10/30 = 0.33; F1 = 2(0.875)(0.824) / (0.875 + 0.824) = 0.848. The model finds cats well but is poor at recognizing non-cats: it calls two thirds of them cats.
Example after K. Markham, “Simple guide to confusion matrix terminology”, Data School.
165 patients are tested. The classifier says “yes” 110 times and “no” 55 times. In reality 105 have the disease and 60 do not. From these totals: TP = 100, FN = 5, FP = 10, TN = 50.
| Actual \ Predicted | yes | no | Total |
|---|---|---|---|
| yes | TP = 100 | FN = 5 | P = 105 |
| no | FP = 10 | TN = 50 | N = 60 |
| Total | P' = 110 | N' = 55 | 165 |
Precision = 100/110 = 0.909; recall = TP / P = 100/105 = 0.952; accuracy = 150/165 = 0.909; error rate = 15/165 = 0.091; specificity = TN / N = 50/60 = 0.833; F1 = 2(0.909)(0.952) / (0.909 + 0.952) = 0.930.
Watch the denominators: recall divides by the actual positives (105), not the predicted ones (110); specificity divides by the actual negatives (60), not the predicted ones (55).
With m classes the confusion matrix is m × m. TP, FP, FN and TN are no longer single numbers: we compute them for each class, treating that class as positive and all the others as negative (one-vs-rest).
| Actual \ Predicted | Class 0 | Class 1 | Class 2 |
|---|---|---|---|
| Class 0 | 3 | 0 | 0 |
| Class 1 | 7 | 50 | 12 |
| Class 2 | 0 | 0 | 18 |
Class 0: TP = 3, FN = 0, FP = 7, TN = 80 → precision 3/10 = 0.30, recall 3/3 = 1.00. Class 1: TP = 50, FN = 7 + 12 = 19, FP = 0, TN = 21 → precision 1.00, recall 50/69 = 0.725. Class 2: TP = 18, FN = 0, FP = 12, TN = 60 → precision 18/30 = 0.60, recall 1.00.
| Method | How | Effect |
|---|---|---|
| Macro average | Compute precision (recall, F1) per class, then take the plain mean | Every class counts equally — small classes matter as much as large ones |
| Micro average | Add up TP, FP, FN over all classes first, then compute one precision/recall | Every example counts equally; for single-label problems it equals accuracy |
| Weighted average | Mean of per-class scores weighted by class size | Between the two |
Three-class example: macro precision = (0.30 + 1.00 + 0.60) / 3 = 0.633; macro recall = (1.00 + 0.725 + 1.00) / 3 = 0.908; micro precision = micro recall = accuracy = (3 + 50 + 18) / 90 = 0.789.
An e-mail classifier, tested against gold labels. (Example from D. Jurafsky & J. H. Martin, Speech and Language Processing, 3rd ed. draft, Ch. 4.)
| Predicted \ Gold | urgent | normal | spam | Precision |
|---|---|---|---|---|
| urgent | 8 | 10 | 1 | 8/19 = 0.42 |
| normal | 5 | 60 | 50 | 60/115 = 0.52 |
| spam | 3 | 30 | 200 | 200/233 = 0.86 |
| Recall | 8/16 = 0.50 | 60/100 = 0.60 | 200/251 = 0.80 |
Note that this table puts the prediction in the rows, so precision reads along a row and recall down a column — always check which way a matrix is laid out. Macro precision = (0.42 + 0.52 + 0.86) / 3 = 0.60; macro recall = (0.50 + 0.60 + 0.80) / 3 = 0.63; accuracy (micro) = 268 / 367 = 0.73. The classifier is good on spam and weak on urgent mail — the macro average exposes that, the micro average hides it.
Precision, recall and the rates all depend on the decision threshold. The ROC (Receiver Operating Characteristic) curve shows the performance at every threshold at once: it plots the true positive rate against the false positive rate. The name comes from radar operators in World War II, who had to tell enemy aircraft from noise.
Put the 14 points on a line by their score and slide the boundary from one end to the other. At each position count TPR = TP / 7 and FPR = FP / 7 and plot the point (FPR, TPR). With everything classified positive we are at (1, 1); with nothing positive at (0, 0); the positions in between trace the curve.
Join the points with straight lines and add the trapezoids: 0.143 × (0.429 + 0.571)/2 + 0.143 × (0.571 + 0.857)/2 + 0.285 × (0.857 + 1)/2 + 0.429 × 1 = 0.072 + 0.102 + 0.265 + 0.429 = 0.87.
The AUC (Area Under the Curve) summarizes the whole curve in one number from 0 to 1. It equals the probability that the model gives a randomly chosen positive a higher score than a randomly chosen negative.
| AUC | Meaning |
|---|---|
| 1.0 | Perfect separation of the classes |
| 0.9 – 1.0 | Excellent |
| 0.8 – 0.9 | Good |
| 0.7 – 0.8 | Fair |
| 0.5 | No better than random guessing |
| < 0.5 | Worse than random — the predictions are inverted |
With k-fold cross-validation you get one ROC curve per fold. Plot them together with their mean curve and report the AUC as mean ± standard deviation — the spread shows how stable the model is.
For regression there is no confusion matrix: we measure how far the predictions are from the true values. Draw a line through the points — but is it a good line or a bad one?
| Metric | Units | Strength | Weakness |
|---|---|---|---|
| MAE | Same as y | Easy to read; robust to outliers | Not differentiable at 0, so awkward for gradient descent |
| MSE | y squared | Smooth and differentiable — the standard training loss | Large errors dominate; hard-to-read units |
| RMSE | Same as y | “Typical error” on the original scale | Still sensitive to outliers |
What is the simplest possible model? A horizontal line at the mean of y. R² compares our model's error with that simple model's error:
Adjusted R² = 1 − (1 − R²)(n − 1)/(n − k − 1) penalizes extra features (Module 3, Section 7).
For the five points of Module 3 (from Statistics How To; ŷ = 9.2 + 0.8x): errors −2.6, 0.6, 3.8, 1.0, −2.8. MAE = 10.8 / 5 = 2.16; MSE = 30.4 / 5 = 6.08; RMSE = 2.47; R² = 1 − 30.4 / 36.8 = 0.17. The single error of 3.8 contributes 47% of the MSE but only 35% of the MAE — squaring magnifies large errors.
| Situation | Use | Why |
|---|---|---|
| Balanced classes, equal costs | Accuracy | Simple and fair when nothing is rare |
| Missing a positive is costly (disease, fraud) | Recall (sensitivity), F2 | Count the misses |
| A false alarm is costly (spam, legal decisions) | Precision, F0.5 | Count the false alarms |
| Imbalanced classes | F1, precision–recall curve, balanced accuracy | Accuracy is misleading |
| Comparing models across all thresholds | ROC curve, AUC | Threshold-independent |
| Several classes of unequal size | Macro-averaged F1 | Small classes are not drowned out |
| Regression with outliers | MAE | Not dominated by a few large errors |
| Regression where large errors are dangerous | RMSE | Punishes large errors |
| How much of the variation is explained | R² / adjusted R² | Unit-free, compared with the mean model |
Run this in Google Colab. It reproduces Example 1, the 4-class matrix from Exercise 4, and the regression example, then evaluates a classifier with stratified 10-fold cross-validation on five metrics.
In part 2, which class has the worst precision, and why? (Look at column a.) In part 4, add "balanced_accuracy" to the list of metrics.
The logistic-regression notebook from class also computes all these metrics: Logistic regression example (.ipynb)
Both books are free to read online from their authors: statlearning.com (T1) and probml.github.io (T2).
100 people: 40 are pregnant, of whom 30 are classified correctly; of the 60 who are not pregnant, 55 are classified correctly. Build the confusion matrix and compute accuracy, precision, recall and F1.
TP = 30, FN = 10, FP = 5, TN = 55. Accuracy = 85/100 = 0.85; precision = 30/35 = 0.857; recall = 30/40 = 0.75; F1 = 2(0.857)(0.75) / (0.857 + 0.75) = 0.80.
For the cat detector of Example 1, compute TPR, FNR, TNR and FPR, and check that each pair adds to 1.
TPR = 140/170 = 0.824, FNR = 30/170 = 0.176 (sum 1). TNR = 10/30 = 0.333, FPR = 20/30 = 0.667 (sum 1). A false-positive rate of 67% is the detector's real weakness.
Model A: precision 0.9, recall 0.1. Model B: precision 0.6, recall 0.5. Compute the arithmetic mean and F1 of each. Which model do the two measures prefer? Then compute F2 for model B.
Means: A = 0.50, B = 0.55. F1: A = 2(0.9)(0.1)/1.0 = 0.18; B = 2(0.6)(0.5)/1.1 = 0.545. The mean barely separates them, but F1 shows that A (which finds only 10% of the positives) is far worse. F2 for B = 5(0.6)(0.5) / (4(0.6) + 0.5) = 1.5 / 2.9 = 0.517.
Rows are the actual class, columns the predicted class:
| Actual \ Predicted | a | b | c | d | Total |
|---|---|---|---|---|---|
| a | 5 | 23 | 17 | 17 | 62 |
| b | 10 | 540 | 21 | 14 | 585 |
| c | 166 | 96 | 436 | 110 | 808 |
| d | 1 | 2 | 5 | 87 | 95 |
| Total | 182 | 661 | 479 | 228 | 1550 |
Compute precision and recall for each class, the macro averages, and the accuracy.
Precision (diagonal / column total): a = 5/182 = 0.027, b = 540/661 = 0.817, c = 436/479 = 0.910, d = 87/228 = 0.382. Recall (diagonal / row total): a = 5/62 = 0.081, b = 540/585 = 0.923, c = 436/808 = 0.540, d = 87/95 = 0.916. Macro precision = 0.534, macro recall = 0.615. Accuracy = (5 + 540 + 436 + 87) / 1550 = 1068/1550 = 0.689. Class a is almost never recognized, and 166 class-c examples are wrongly called a — that is where to improve.
Six test examples with scores and labels: 0.9 (+), 0.8 (+), 0.7 (−), 0.6 (+), 0.4 (−), 0.2 (−). Compute (FPR, TPR) at thresholds 0.75 and 0.5. Then compute the AUC as the fraction of (positive, negative) pairs in which the positive has the higher score.
Threshold 0.75: positives = {0.9, 0.8} → TP = 2, FP = 0 → (FPR, TPR) = (0, 0.667). Threshold 0.5: positives = {0.9, 0.8, 0.7, 0.6} → TP = 3, FP = 1 → (0.333, 1). AUC: 3 × 3 = 9 pairs; 0.9 and 0.8 beat all three negatives (6 pairs), 0.6 beats 0.4 and 0.2 (2 pairs) → AUC = 8/9 = 0.89.
Which metric would you report for: (a) a fraud detector where 0.2% of transactions are fraud; (b) a house-price model where a few luxury homes have huge prices; (c) comparing two credit-scoring models before the bank has chosen a cut-off?
(a) Recall and precision (or F1 / the precision–recall curve) on the fraud class — accuracy would be 99.8% for a model that flags nothing. (b) MAE, so the few luxury homes do not dominate (report RMSE too if large errors are expensive). (c) ROC curve and AUC — they compare the models over all possible cut-offs.
We now have the tools to compare classifiers. Module 6 introduces a classifier built from rules instead of equations: the decision tree.