Fit the straight line (or plane) that best predicts a number — by least squares in one step, or by gradient descent one step at a time — and keep it from overfitting.
Module 3 · Week 4 · Lecture notes by Dr. Abdulkarim Albanna
Core Supervised · Regression ~70 miny = b0 + b1x + εPrerequisites: Modules 1–2, plus basic algebra and the idea of a derivative.
Regression is supervised learning where the target is a real number. It models how a dependent variable (the target y, the thing we predict) changes with one or more independent variables (the features x). The model is trained on data with known answers so that it can predict the answer for new, unseen examples.
| Application | Features x | Target y |
|---|---|---|
| House valuation | Size, rooms, location, age | Sale price |
| Finance | Past prices, trading volume | Tomorrow's stock price |
| Lending | Income, debts, payment history | Credit score |
| Marketing | TV, radio and newspaper budgets | Sales |
| Education | Hours studied | Exam score |
We assume the target is a straight-line function of the feature plus random noise:
b0 is the intercept (the value of y when x = 0), b1 is the slope (how much y changes when x grows by one unit), and εi is the random error — everything the line cannot explain. In machine-learning notation the same model is written ŷ = w·x + b, with w the weight and b the bias.
Linear regression learns to predict a real-valued target as a linear function of the inputs. Like every algorithm in this course, it follows three steps: choose a model, define a loss that says how bad the fit is, and fit the model with an optimization algorithm.
The loss for one example is the squared error between the prediction ŷ and the true target t. Squaring makes every error positive and punishes large errors much more than small ones. The cost averages the loss over all N training examples:
The ½ is only for convenience — it cancels when we differentiate. Training means finding the w and b that make J as small as possible.
The cost is a smooth bowl, so its minimum is where the partial derivatives with respect to both parameters are zero. Setting ∂J/∂b1 = 0 and ∂J/∂b0 = 0 and solving gives the least-squares estimates:
yi − (m xi + b) — positive above the line, negative below. Squaring and adding them gives SEline. (Redrawn from Dr. Albanna’s slides, which follow Khan Academy’s proof of the least-squares formulas.)Write the line as y = m x + b (slope m, intercept b). The squared error of the line over n points is
Since (y1² + … + yn²) / n is the mean of y², the sum equals n·mean(y²); the same holds for the other sums:
As a function of m and b, SEline is a bowl: m² and b² appear with positive coefficients. Its lowest point is where the surface is flat in both directions — where both partial derivatives are zero:
∂SE/∂m = 0 and ∂SE/∂b = 0. (Redrawn from Dr. Albanna’s slides, which follow Khan Academy’s proof of the least-squares formulas.)Divide the second equation by 2n: b = ȳ − m·x̄ — so the line passes through (x̄, ȳ). Substitute this b into the first equation and solve for m:
These are exactly the formulas above (m = b1, b = b0).
From the slide's sums (n = 7): Σx = 5.15, Σy = 71.403, Σxy = 49.458, Σx² = 3.9869. Means: x̄ = 0.7357, ȳ = 10.2004, mean(xy) = 7.0654, mean(x²) = 0.5696.
m = (7.0654 − 0.7357 × 10.2004) / (0.5696 − 0.7357²) = (7.0654 − 7.5046) / (0.5696 − 0.5413) = −0.4392 / 0.0283 = −15.53 — the same slope as the table method, up to the rounding of the sums.
The fitted line always passes through the point of means (x̄, ȳ). An equivalent form that avoids subtracting the means is b1 = (nΣxy − ΣxΣy) / (nΣx² − (Σx)²).
Seven observations of an input H (x) and a target T (y). Step 1: the means are H̄ = 5.15 / 7 = 0.7357 and T̄ = 71.403 / 7 = 10.2004. Step 2: fill in the deviation columns.
| H (x) | T (y) | x − x̄ | y − ȳ | (x − x̄)(y − ȳ) | (x − x̄)² |
|---|---|---|---|---|---|
| 0.89 | 7.330 | 0.1543 | −2.8704 | −0.4429 | 0.0238 |
| 0.85 | 7.223 | 0.1143 | −2.9774 | −0.3403 | 0.0131 |
| 0.89 | 8.440 | 0.1543 | −1.7604 | −0.2716 | 0.0238 |
| 0.56 | 12.230 | −0.1757 | 2.0296 | −0.3566 | 0.0309 |
| 0.84 | 8.340 | 0.1043 | −1.8604 | −0.1940 | 0.0109 |
| 0.69 | 13.300 | −0.0457 | 3.0996 | −0.1417 | 0.0021 |
| 0.43 | 14.540 | −0.3057 | 4.3396 | −1.3267 | 0.0935 |
| Sum | −3.0738 | 0.1980 | |||
Step 3: b1 = −3.0738 / 0.1980 = −15.526. Step 4: b0 = 10.2004 − (−15.526)(0.7357) = 21.623.
Model: T = −15.526·H + 21.623. For a new input H = 0.70 it predicts T = −15.526(0.70) + 21.623 = 10.75.
Least squares has a closed-form answer only because the model is so simple. Gradient descent is a second way to minimize the cost that works for almost every model in machine learning, including neural networks. It is iterative: start from any weights (e.g. all zeros) and repeatedly step in the direction of steepest descent — opposite to the gradient — until the cost stops falling.
η (eta) is the learning rate, the size of each step. Both parameters are updated at the same time, using the gradient computed with the old values.
Data: x = [1, 2, 3], t = [2, 4, 6] (the true answer is w = 2, b = 0). Start at w = 0, b = 0 with η = 0.1.
Predict: ŷ = [0, 0, 0]; errors ŷ − t = [−2, −4, −6]. Cost J = (4 + 16 + 36) / (2·3) = 9.33.
Gradients: ∂J/∂w = (−2·1 − 4·2 − 6·3) / 3 = −28/3 = −9.33; ∂J/∂b = (−2 − 4 − 6) / 3 = −4.
Update: w = 0 − 0.1(−9.33) = 0.933; b = 0 − 0.1(−4) = 0.4. New cost J = 1.88 — down from 9.33. One more step gives w = 1.351, b = 0.573, J = 0.40; after a few hundred steps w → 2, b → 0.
| Least squares (closed form) | Gradient descent | |
|---|---|---|
| How | Solve ∂J = 0 directly | Repeat small downhill steps |
| Hyperparameters | None | Learning rate, number of iterations |
| Large data / many features | Slow (matrix inverse) | Scales well (and mini-batch versions scale further) |
| Other models | Only linear regression | Logistic regression, neural networks, … |
When features have very different ranges, the cost bowl becomes a long narrow valley and gradient descent zig-zags. Standardizing the features first (Module 2) makes it converge much faster.
With p features the line becomes a plane (or a hyperplane in more dimensions):
Each coefficient wj is the change in the prediction when xj grows by one unit while all other features stay fixed. Stacking the N training examples as the rows of a matrix X (with a column of 1s for the bias), least squares has a matrix solution, the normal equation:
A house-price model: price = 50 + 120·size + 15·rooms − 2·age (price in $1000s, size in 100 m², age in years). A 150 m² house with 4 rooms, 10 years old: 50 + 120(1.5) + 15(4) − 2(10) = 50 + 180 + 60 − 20 = 270 → $270,000. Reading the coefficients: one extra room adds $15,000 for a house of the same size and age; each year of age subtracts $2,000.
Three families of metrics evaluate a regression model (Module 5 covers them in full):
Split each point's distance from the mean into a part the line explains and a part it does not. Summed over all points, SStotal = SSregression + SSresidual, and
R² runs from 0 (the line is no better than predicting the mean) to 1 (a perfect fit). It is positive even when the slope is negative — it measures fit, not direction.
Plain R² never goes down when you add a feature, even a useless one. Adjusted R² charges a penalty for every extra feature (n = observations, k = features):
Example from Statistics How To, “Mean Squared Error”.
Points (43, 41), (44, 45), (45, 49), (46, 47), (47, 44); least squares gives ŷ = 9.2 + 0.8x.
| x | y | ŷ | error y − ŷ | error² | (y − ȳ)² |
|---|---|---|---|---|---|
| 43 | 41 | 43.6 | −2.6 | 6.76 | 17.64 |
| 44 | 45 | 44.4 | 0.6 | 0.36 | 0.04 |
| 45 | 49 | 45.2 | 3.8 | 14.44 | 14.44 |
| 46 | 47 | 46.0 | 1.0 | 1.00 | 3.24 |
| 47 | 44 | 46.8 | −2.8 | 7.84 | 1.44 |
| Sum (ȳ = 45.2) | 30.40 | 36.80 | |||
MSE = 30.4 / 5 = 6.08; RMSE = √6.08 = 2.47; MAE = (2.6 + 0.6 + 3.8 + 1.0 + 2.8) / 5 = 2.16; R² = 1 − 30.4 / 36.8 = 0.174; R²adj = 1 − (0.826)(4 / 3) = −0.10. A negative adjusted R² is a warning: with only 5 points, x explains almost nothing about y.
The least-squares estimates are trustworthy only when the data meets a few assumptions. Check them with a residual plot (residual against prediction):
| Assumption | Meaning | Symptom when it fails |
|---|---|---|
| Linearity | The relationship between x and y is a straight line | Residuals form a curve |
| Independence | Observations (and their errors) do not depend on each other | Patterns over time in the residuals (autocorrelation) |
| Normality | Errors are normally distributed around the line | Skewed histogram of residuals |
| Equal variance (homoscedasticity) | The spread of y is the same at every x | Residuals fan out in a funnel |
| No multicollinearity (multiple regression) | Features are not strongly correlated with each other | Unstable coefficients with unexpected signs |
When the relationship is curved, we can still use linear regression: create new features x², x³, … and fit a linear model on them. The model stays linear in its weights, so least squares and gradient descent work unchanged:
Regularization adds a penalty on the size of the weights to the cost. The hyperparameter α (also written λ) sets how strong the penalty is.
| Ridge (L2) | Lasso (L1) | |
|---|---|---|
| Penalty | Sum of squared weights | Sum of absolute weights |
| Effect | Shrinks all weights toward 0, but none to exactly 0 | Shrinks and sets some weights exactly to 0 |
| Use it when | Many features all contribute a little; multicollinearity | You suspect only a few features matter — it does feature selection |
β1 = 0; the Ridge circle (right) has no corners. (Textbook T1: James et al., ISLP, Figure 6.7.)α = 0 gives ordinary least squares (risk of overfitting); a very large α pushes all weights to 0 (underfitting). Split the training data into train and validation sets, train one model per candidate α (e.g. 0.001, 0.01, 0.1, 1, 10, 100), and keep the one with the lowest validation error — or use cross-validation. Report the final score on the untouched test set. Standardize the features first, so the penalty treats them all equally.
Weights w = (3, −2, 0.5), sum of squared errors 40, α = 2. Ridge: 40 + 2(9 + 4 + 0.25) = 40 + 26.5 = 66.5. Lasso: 40 + 2(3 + 2 + 0.5) = 40 + 11 = 51. Training then trades a little extra error for smaller weights.
Run this in Google Colab. Part 1 reproduces the worked example three ways — by the formula, with scikit-learn, and with gradient descent written from scratch — and all three give the same line. Part 2 fits multiple, Ridge and Lasso regression on the built-in diabetes dataset (10 features).
Lasso with α = 1 kept only 2 of the 10 features — too strong a penalty. Try α = 0.01, 0.1, 0.5: how many coefficients are zero each time, and which α gives the best test R²? In part 1, set eta = 5 and watch gradient descent diverge.
More examples from class (house size vs. price, California housing, study hours vs. exam score) are in the notebooks: LR Part 1 (.ipynb) LR Part 2 (.ipynb)
Both books are free to read online from their authors: statlearning.com (T1) and probml.github.io (T2).
X = GDP, Y = passenger-vehicle sales (in 0.1 millions): 2011 (6.2, 26.3), 2012 (6.5, 26.65), 2013 (5.48, 25.03), 2014 (6.54, 26.01), 2015 (7.18, 27.9), 2016 (7.93, 30.47). Find (1) the regression equation, (2) MSE and RMSE, (3) MAE, (4) R², (5) the predicted sales when GDP = 7.23.
Means: x̄ = 6.638, ȳ = 27.06. Σ(x−x̄)(y−ȳ) = 7.704, Σ(x−x̄)² = 3.524. (1) b1 = 7.704 / 3.524 = 2.186, b0 = 27.06 − 2.186(6.638) = 12.549 → Y = 12.549 + 2.186X. Predictions: 26.10, 26.76, 24.53, 26.85, 28.24, 29.88; errors: 0.198, −0.108, 0.502, −0.835, −0.344, 0.587. (2) MSE = 0.244, RMSE = 0.494. (3) MAE = 0.429. (4) R² = 0.920 — GDP explains 92% of the variation in sales. (5) 12.549 + 2.186(7.23) = 28.35, i.e. about 2.84 million vehicles.
Hours studied x = [1, 2, 3, 4, 5], exam score y = [52, 55, 61, 64, 68]. Find the least-squares line and predict the score for 6 hours.
x̄ = 3, ȳ = 60. Deviations x: −2, −1, 0, 1, 2; y: −8, −5, 1, 4, 8. Σ(x−x̄)(y−ȳ) = 16 + 5 + 0 + 4 + 16 = 41; Σ(x−x̄)² = 10. b1 = 4.1, b0 = 60 − 4.1(3) = 47.7. Line: ŷ = 47.7 + 4.1x; for 6 hours: 47.7 + 24.6 = 72.3.
Data x = [1, 2], t = [3, 5]. Start at w = 1, b = 0 with η = 0.1. Compute the cost (with ½N), both gradients, the updated w and b, and the new cost.
ŷ = [1, 2]; errors ŷ − t = [−2, −3]. J = (4 + 9) / 4 = 3.25. ∂J/∂w = (−2·1 − 3·2) / 2 = −4; ∂J/∂b = (−2 − 3) / 2 = −2.5. Update: w = 1 + 0.4 = 1.4, b = 0 + 0.25 = 0.25. New ŷ = [1.65, 3.05]; errors [−1.35, −1.95]; J = (1.8225 + 3.8025) / 4 = 1.41 — the cost fell from 3.25 to 1.41.
Model A uses 2 features and has R² = 0.80; model B adds 8 more features and reaches R² = 0.82. Both were fit on n = 30 observations. Compute adjusted R² for each. Which model do you prefer?
A: 1 − 0.20 × 29/27 = 1 − 0.2148 = 0.785. B: 1 − 0.18 × 29/19 = 1 − 0.2747 = 0.725. Prefer A: the 8 extra features raised R² only slightly, and adjusted R² shows they do not pay for their complexity.
After fitting income against years of experience, the residuals are small for junior staff and grow steadily larger for senior staff. Which assumption is violated, and what could you do?
The residuals fan out like a funnel: equal variance (homoscedasticity) is violated. Remedies: transform the target (e.g. model log(income)), use weighted least squares, or report robust standard errors.
(a) You have 500 genes and believe only about 10 affect a disease. (b) You have 20 economic indicators, many strongly correlated, all believed to matter a little. (c) What happens to both models as α → 0, and as α → ∞?
(a) Lasso — it sets the irrelevant genes' weights to exactly 0, selecting the few that matter. (b) Ridge — it shrinks correlated weights together and keeps them all. (c) As α → 0 both become ordinary least squares (may overfit); as α → ∞ all weights go to 0 and the model predicts a constant (underfits).
ŷ = w·x + b; training minimizes the mean squared error.b1 = Σ(x−x̄)(y−ȳ) / Σ(x−x̄)², b0 = ȳ − b1x̄; the normal equation does the same for many features.w ← w − η·∂J/∂w, and works for models with no closed form.Regression predicts numbers. What if the answer is yes or no? Module 4 bends the straight line through a sigmoid to get logistic regression, our first classifier.