All Modules Sigmoid Odds & Logit Training Threshold Exercises

Logistic Regression

Our first classifier: bend a straight line through the sigmoid to get the probability of a yes/no outcome, train it by maximum likelihood, and choose where to draw the line between yes and no.

Module 4 · Weeks 4–5 · Lecture notes by Dr. Abdulkarim Albanna

Core Supervised · Classification ~65 min

What You'll Learn

  • What a binary classification problem is, and why linear regression is the wrong tool for it
  • The sigmoid (logistic) function and the logistic model p = σ(w·x + b)
  • Odds, log-odds (logit) and how to read a coefficient as an odds ratio
  • Training by maximum likelihood: the log-likelihood, the log-loss and gradient descent
  • The decision threshold and the decision boundary, and how moving the threshold trades one error for another
  • Multiple logistic regression, worked on a real medical dataset

Prerequisite: Module 3. Logistic regression reuses the linear model w·x + b and gradient descent from linear regression.

1. From Regression to Classification

Often we care about an input–output relationship, just as in regression, but the output is discrete, not continuous: spam or not, sick or healthy, pass or fail. Guessing the class from the inputs is classification. In binary classification the target is t ∈ {0, 1}; examples with t = 1 are the positive examples and those with t = 0 the negative ones.

Rather than guessing the class directly, logistic regression models the probability P(Y = 1 | X = x) as a function of x, and estimates the unknown parameters by maximum likelihood. Despite its name it is a classification method; “regression” refers to the linear model it fits inside.

Running example: age and coronary disease

33 adults, their age, and whether they show signs of coronary heart disease (CD = 1) or not (CD = 0). 14 have the disease. The mean age is 38.6 years in the healthy group and about 59 years in the diseased group — so age matters. Grouping by age makes the pattern clear:

Age group20–2930–3940–4950–5960–6970–7980–89
People in group5677521
Diseased0124421
% diseased0%17%29%57%80%100%100%

The percentage climbs along an S-shape from 0 to 100%. That S is what logistic regression fits.

Left: a straight line through the 0/1 disease data goes below 0 and above 1. Right: an S-shaped logistic curve stays between 0 and 1 and follows the age-group percentages.
Left: a straight line through 0/1 data predicts “probabilities” below 0 for the young and above 1 for the old — meaningless. Right: the logistic curve stays inside [0, 1] and tracks the percentage diseased in each age group (green squares).
Linear regressionLogistic regression
PredictsA continuous valueThe probability of a class
ShapeStraight lineS-shaped (sigmoid) curve
TaskRegressionClassification
Fitted byLeast squares / gradient descentMaximum likelihood (gradient descent)
Fit measured byMSE, R²Log-loss (−2 log-likelihood), accuracy, precision, recall

2. Types of Logistic Regression

TypeTargetExamples
Binomial (binary)Two categoriesYes / No, Pass / Fail, Spam / Not spam — the most common, and the focus of this module
MultinomialThree or more unordered categoriesCat / Dog / Sheep; which product a customer buys (uses softmax)
OrdinalThree or more ordered categoriesLow / Medium / High; a 1–5 star rating

3. The Sigmoid Function

The sigmoid (or logistic) function takes any real number and squashes it into the range (0, 1), drawing an S-shaped curve:

\[ \sigma(z) = \frac{1}{1 + e^{-z}} \]
  • σ(z) → 1 as z → +∞, and σ(z) → 0 as z → −∞; it is always strictly between 0 and 1.
  • σ(0) = 0.5; the curve is symmetric: σ(−z) = 1 − σ(z).
  • Its derivative is simple: σ'(z) = σ(z)(1 − σ(z)) — handy for gradient descent.
  • Scaling and shifting the input changes the curve: σ(5x) is steeper, and σ(5x − 10) is steeper and centred at x = 2.
The sigmoid curve with sigma(-2)=0.12, sigma(0)=0.5 and sigma(2)=0.88 marked, and the regions z below and above 0
The sigmoid. Positive scores give probabilities above 0.5, negative scores below 0.5.

4. The Logistic Regression Model

Logistic regression takes the linear model from Module 3, z = w·x + b = w1x1 + … + wnxn + b, whose output is any real number, and feeds it through the sigmoid to get a probability:

\[ P(y = 1 \mid \mathbf{x}) = \sigma(\mathbf{w}^\top\mathbf{x} + b) = \frac{1}{1 + e^{-(\mathbf{w}^\top\mathbf{x} + b)}} \qquad P(y = 0 \mid \mathbf{x}) = 1 - P(y = 1 \mid \mathbf{x}) \]
xfeatures Linear scorez = w·x + b Sigmoidσ(z) = 1 / (1 + e⁻ᶻ) Probabilityp = P(y = 1 | x) Thresholdŷ = 1 if p ≥ 0.5 age 60 z = 1.197 σ(1.197) = 0.768 p = 0.768 0.768 ≥ 0.5 → ŷ = 1
How logistic regression predicts, traced for a 60-year-old in the coronary-disease example: a linear score, squashed by the sigmoid into a probability, then cut by a threshold into a class.
Weight on the x-axis, an S-curve from not obese to obese, and arrows reading the probability for three new mice
Reading the curve: a heavy mouse has a high probability of being obese, a middle one about 50%, a light one a small probability. Although the output is a probability, it is usually used to classify. (From Dr. Albanna's slides; image from StatQuest with Josh Starmer, “Logistic Regression”.)

5. Odds, Log-Odds and the Logit

Why the sigmoid? Because it is the inverse of the logit, the function that turns a probability into a number a straight line can model.

  • Odds = the probability that an event happens divided by the probability that it does not: odds = p / (1 − p). If P(Y = 1) = 0.8, the odds are 0.8 / 0.2 = 4 (“4 to 1”).
  • Log-odds or logit = the natural logarithm of the odds: logit(p) = log(p / (1 − p)). It maps a probability in (0, 1) onto the whole real line (−∞, +∞).
probability pbetween 0 and 1 0 0.25 0.5 0.75 1 0.1 0.5 0.8 0.9 odds p / (1 − p)between 0 and ∞ 0 1 2 4 6 8 10 0.11 1 4 9 log-odds log(p / (1 − p))between −∞ and +∞ -3 -2 -1 0 1 2 3 -2.20 0 +1.39 +2.20
One probability, three scales. Odds stretch [0, 1] onto [0, ∞) but crowd everything below p = 0.5 between 0 and 1; log-odds spread it symmetrically over the whole number line, with p = 0.5 at 0. That symmetric, unbounded scale is what a straight line w·x + b can model.

Logistic regression assumes the log-odds are linear in x. Solving for p gives back the sigmoid:

\[ \log\frac{p}{1-p} = \mathbf{w}^\top\mathbf{x} + b \quad\Longleftrightarrow\quad \frac{p}{1-p} = e^{\mathbf{w}^\top\mathbf{x} + b} \quad\Longleftrightarrow\quad p = \frac{1}{1 + e^{-(\mathbf{w}^\top\mathbf{x} + b)}} \]

Reading the coefficients

  • b is the log-odds when all features are 0.
  • wj is the change in log-odds when xj grows by one unit (all else fixed).
  • ewj is the odds ratio: the odds are multiplied by ewj for each one-unit increase in xj.

Worked example — survival by age

A model of survival for patients exposed to a disease: log(P / (1 − P)) = 1.8158 − 0.0665·Age.

Newborn (Age = 0): log-odds = 1.8158 → odds = e1.8158 = 6.15 → P = 6.15 / (1 + 6.15) = 0.86.

Age 25: log-odds = 1.8158 − 0.0665(25) = 0.1533 → odds = e0.1533 = 1.17 → P = 1.17 / 2.17 = 0.54.

The odds ratio per year is e−0.0665 = 0.936: each extra year multiplies the odds of survival by 0.936, about a 6.4% drop.

6. Training: Maximum Likelihood

Logistic regression has no residuals in the least-squares sense, so we cannot use least squares or R². Instead it uses maximum likelihood estimation (MLE): choose the parameters that make the observed labels as probable as possible.

Likelihood and log-likelihood

For each example the model gives pi = P(yi = 1 | xi). The probability it assigns to the label actually observed is pi if yi = 1 and 1 − pi if yi = 0. Assuming independent examples, the likelihood of the whole dataset is the product, and since log is increasing we maximize the easier log-likelihood (a sum):

\[ L(\mathbf{w}, b) = \prod_{i=1}^{n} p_i^{\,y_i}(1-p_i)^{1-y_i} \qquad\quad \ell(\mathbf{w}, b) = \sum_{i=1}^{n}\big[y_i \log p_i + (1-y_i)\log(1-p_i)\big] \]

Because every pi < 1, ℓ is negative. Maximizing ℓ is the same as minimizing −2ℓ (the deviance) or the average log-loss (binary cross-entropy):

\[ J(\mathbf{w}, b) = -\frac{1}{n}\sum_{i=1}^{n}\big[y_i \log p_i + (1-y_i)\log(1-p_i)\big] \]
Loss minus log p for a positive example and minus log(1-p) for a negative example
The log-loss for one example. If the truth is y = 1, predicting p = 0.9 costs only 0.11, p = 0.5 costs 0.69, and a confident wrong p = 0.1 costs 2.30 — confident mistakes are punished hardest.
Several candidate S-curves; the one with the maximum likelihood is selected
MLE in pictures: try a curve, compute the likelihood of all the points, shift the curve and compute it again; the curve with the maximum likelihood wins. (From Dr. Albanna's slides; image from StatQuest with Josh Starmer, “Logistic Regression”.)

Gradient descent

There is no closed-form solution, so the parameters are found iteratively: start from zeros, compute the log-likelihood, adjust the coefficients, and repeat until it stops improving. The gradient of the log-loss has the same shape as in linear regression — only ŷ is now a sigmoid:

\[ \frac{\partial J}{\partial w_j} = \frac{1}{n}\sum_i (p_i - y_i)\,x_{ij} \qquad \frac{\partial J}{\partial b} = \frac{1}{n}\sum_i (p_i - y_i) \qquad w_j \leftarrow w_j - \eta\,\frac{\partial J}{\partial w_j} \]

At the optimum, Σ(yi − pi)xij = 0 for every feature — the condition the maximum-likelihood estimates solve.

Worked example — one gradient step

Data x = [1, 2], y = [0, 1]; start at w = 0, b = 0, η = 1. Then z = [0, 0], p = [0.5, 0.5], and the log-loss is −(log 0.5 + log 0.5) / 2 = 0.693.

Gradients: ∂J/∂w = ((0.5 − 0)·1 + (0.5 − 1)·2) / 2 = −0.25; ∂J/∂b = ((0.5 − 0) + (0.5 − 1)) / 2 = 0. Update: w = 0.25, b = 0.

Now p = [σ(0.25), σ(0.5)] = [0.562, 0.622] and the log-loss is −(log 0.438 + log 0.622) / 2 = 0.650 — lower than 0.693.

Worked example — MLE by hand

Four observations from a binomial distribution with n = 3 trials each: (1, 3, 2, 2) successes. Find the success probability θ.

Likelihood: L(θ) = C(3,1)θ(1−θ)² · C(3,3)θ³ · C(3,2)θ²(1−θ) · C(3,2)θ²(1−θ) = 27θ8(1−θ)4.

Log-likelihood: ℓ(θ) = ln 27 + 8 ln θ + 4 ln(1−θ). Derivative: 8/θ − 4/(1−θ) = 0 → 8(1−θ) = 4θ → θ̂ = 8/12 = 2/3 — total successes over total trials, as intuition says.

7. The Decision Threshold and the Decision Boundary

To turn a probability into a class we pick a threshold t: predict class 1 if p ≥ t, else class 0. The default is t = 0.5. Since σ(z) ≥ 0.5 exactly when z ≥ 0, the default rule is simply w·x + b ≥ 0. The set of points where w·x + b = 0 is the decision boundary — a point for one feature, a line for two, a plane for three: logistic regression is a linear classifier.

Left: a plane with its normal vector w. Right: iris petal length vs width with a straight dashed boundary separating Virginica from the rest
(a) The boundary w·x + b = 0 is a plane whose normal is w. (b) A logistic regression on two iris features: the dashed line separates Iris-Virginica from the other flowers. (Textbook T2: K. P. Murphy, Probabilistic Machine Learning: An Introduction, MIT Press 2022, Figure 10.1; CC BY-NC-ND.)

Moving the threshold trades one kind of error for another. A low threshold flags more positives: it catches more true cases (higher recall) but raises more false alarms. A high threshold is more cautious: fewer false alarms (higher precision) but more missed cases. In medicine, missing a sick patient is usually worse, so the threshold is set below 0.5. Module 5 measures these trade-offs with precision, recall and the ROC curve.

The fitted disease curve with thresholds 0.3, 0.5 and 0.7 and the matching cut-off ages 45, 51 and 57
The coronary-disease model with three thresholds. Each threshold becomes an age cut-off: predict disease from age 45 (t = 0.3), 51 (t = 0.5) or 57 (t = 0.7).
ThresholdCut-off ageTrue positivesFalse positivesMissed (false negatives)True negativesAccuracy
0.3≥ 4512821169.7%
0.5≥ 519351675.8%
0.7≥ 578161878.8%

The threshold with the best accuracy (0.7) misses 6 of the 14 sick patients; the 0.3 threshold misses only 2. Which one is “best” depends on the cost of each kind of mistake, not on accuracy alone.

8. Worked Example: Coronary Disease and Age

Fitting logistic regression by maximum likelihood to the 33 patients from Section 1 gives:

\[ \log\frac{p}{1-p} = -6.820 + 0.1336 \cdot \text{Age} \]

Using the model

Age 60: z = −6.820 + 0.1336(60) = 1.197 → p = 1 / (1 + e−1.197) = 0.768 → 77% chance of disease; at t = 0.5, predict disease.

Age 40: z = −6.820 + 5.344 = −1.475 → p = 0.186 → predict healthy.

Odds ratio: e0.1336 = 1.143 — every extra year multiplies the odds of disease by 1.143 (+14%); ten years multiply them by e1.336 = 3.80.

Decision boundary: z = 0 → Age = 6.820 / 0.1336 = 51.0 years.

Fit: log-likelihood −14.20, so −2LL = 28.40; average log-loss 0.430.

Age3040506070
log-odds z−2.81−1.48−0.141.202.53
odds ez0.060.230.873.3112.6
P(disease)0.0570.1860.4650.7680.926

9. Multiple Logistic Regression

With several features — continuous, binary, ordinal or (one-hot encoded) nominal — the log-odds are a linear combination of all of them:

\[ \log\frac{p}{1-p} = b + w_1 x_1 + w_2 x_2 + \dots + w_n x_n \]

Each wj is now the increase in log-odds for a one-unit increase in xj with all other features held constant — its association with the outcome, adjusted for the others.

Worked example — adding gender

Extending the survival model with Gender (0 = male, 1 = female): log(P / (1 − P)) = 1.8158 − 0.0665·Age + 1.5972·Gender.

Male model (Gender = 0): 1.8158 − 0.0665·Age. Female model (Gender = 1): (1.8158 + 1.5972) − 0.0665·Age = 3.4130 − 0.0665·Age.

At age 25: male z = 0.153 → P = 0.54; female z = 1.751 → P = 0.85. The odds ratio for being female is e1.5972 = 4.94: at any age, a woman's odds of survival are about 4.9 times a man's.

Python Lab

Run this in Google Colab. It fits the coronary-disease model with scikit-learn, prints the probabilities and the effect of the threshold, then re-fits the same model with gradient descent written from scratch.

import numpy as np from sklearn.linear_model import LogisticRegression from sklearn.metrics import confusion_matrix, accuracy_score, log_loss age = np.array([22, 23, 24, 27, 28, 30, 30, 32, 33, 36, 38, 40, 41, 46, 47, 48, 49, 49, 50, 51, 51, 52, 54, 55, 58, 60, 60, 62, 66, 67, 71, 77, 81], dtype=float) cd = np.array([0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 0, 1, 0, 1, 0, 1, 0, 0, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1]) # 1. Maximum-likelihood fit (penalty=None: plain logistic regression, no regularization) model = LogisticRegression(penalty=None).fit(age.reshape(-1, 1), cd) b, w = model.intercept_[0], model.coef_[0][0] print(f"log-odds = {b:.3f} + {w:.4f} * age (odds ratio per year = {np.exp(w):.3f})") for a in [30, 40, 50, 60, 70]: print(f"age {a}: P(disease) = {model.predict_proba([[a]])[0, 1]:.3f}") # 2. Log-loss and the effect of the threshold p = model.predict_proba(age.reshape(-1, 1))[:, 1] print("log loss:", round(log_loss(cd, p), 3)) for t in [0.3, 0.5, 0.7]: pred = (p >= t).astype(int) tn, fp, fn, tp = confusion_matrix(cd, pred).ravel() print(f"threshold {t}: TP={tp} FP={fp} FN={fn} TN={tn} accuracy={accuracy_score(cd, pred):.3f}") # 3. The same model by gradient descent on the log-loss (standardize age so it converges fast) z_age = (age - age.mean()) / age.std() w_s, b_s, eta = 0.0, 0.0, 0.5 for step in range(5000): p_s = 1 / (1 + np.exp(-(w_s * z_age + b_s))) w_s -= eta * ((p_s - cd) * z_age).mean() b_s -= eta * (p_s - cd).mean() w_gd = w_s / age.std(); b_gd = b_s - w_s * age.mean() / age.std() # back to years print(f"gradient descent: log-odds = {b_gd:.3f} + {w_gd:.4f} * age") # log-odds = -6.820 + 0.1336 * age (both ways); P(60) = 0.768; thresholds as in the table above

Try it

Remove penalty=None. scikit-learn then applies L2 regularization by default (Module 3): how do the coefficient and the probabilities change? Then try thresholds 0.4 and 0.6 and fill in the table.

The class notebook adds a 2-D example with a plotted decision boundary, the sigmoid and log-loss curves, and every classification metric: Logistic regression example (.ipynb)

Textbook Reading

From the course syllabus

  • T1 James et al. — ISLP, Ch. 4 Classification: §4.1–4.2 and §4.3 Logistic Regression (the model, estimating the coefficients, multiple logistic regression).
  • T2 Murphy — Probabilistic Machine Learning, Ch. 10 Logistic Regression (§10.2 binary logistic regression, including the MLE and gradient descent).

Both books are free to read online from their authors: statlearning.com (T1) and probml.github.io (T2).

Exercises

1

Predict with the model

A model has w = 2, b = −3. Compute P(y = 1) for x = 1 and x = 2, and the class at threshold 0.5. Where is the decision boundary?

x = 1: z = 2 − 3 = −1, p = 1 / (1 + e) = 0.269 → class 0. x = 2: z = 4 − 3 = 1, p = 0.731 → class 1. Boundary: 2x − 3 = 0 → x = 1.5.

2

Probability, odds, log-odds

(a) Convert p = 0.75 to odds and log-odds. (b) Convert log-odds −1 to a probability. (c) What probability has log-odds 0?

(a) odds = 0.75 / 0.25 = 3; log-odds = ln 3 = 1.099. (b) p = 1 / (1 + e1) = 0.269. (c) p = 0.5 (odds = 1).

3

Interpret a coefficient

In the coronary-disease model the age coefficient is 0.1336. What is the odds ratio for one year, and for ten years? Explain each in one sentence.

One year: e0.1336 = 1.143 — each extra year multiplies the odds of disease by 1.143 (14% higher). Ten years: e1.336 = 3.80 — someone ten years older has about 3.8 times the odds. Note it is the odds that are multiplied, not the probability.

4

Compute the log-loss

Three examples with true labels y = [1, 0, 0] and predicted probabilities p = [0.9, 0.3, 0.6]. Compute each loss and the average. Which prediction is the worst?

Example 1 (y = 1): −ln 0.9 = 0.105. Example 2 (y = 0): −ln(1 − 0.3) = 0.357. Example 3 (y = 0): −ln(1 − 0.6) = 0.916. Average = 1.378 / 3 = 0.459. The third is worst: it gave 60% to the wrong class.

5

Move the threshold

Eight predicted probabilities with their true labels: 0.92 (1), 0.75 (1), 0.58 (0), 0.45 (1), 0.40 (0), 0.22 (1), 0.15 (0), 0.08 (0). Count TP, FP, FN and TN at threshold 0.5 and at 0.3. Which threshold would you use to screen for a serious disease?

t = 0.5: predicted 1 = {0.92, 0.75, 0.58} → TP = 2, FP = 1; FN = 2 (0.45, 0.22), TN = 3. t = 0.3: 0.45 and 0.40 also become 1 → TP = 3, FP = 2, FN = 1 (0.22), TN = 2. For screening a serious disease use 0.3: it misses fewer sick patients, at the cost of one more false alarm, which a follow-up test can resolve.

6

Maximum likelihood for a coin

A coin lands heads 7 times in 10 tosses. Write the log-likelihood of θ = P(heads), differentiate, and find the maximum-likelihood estimate.

L(θ) = θ7(1 − θ)3, so ℓ(θ) = 7 ln θ + 3 ln(1 − θ). dℓ/dθ = 7/θ − 3/(1 − θ) = 0 → 7(1 − θ) = 3θ → θ̂ = 7/10 = 0.7.

Recap & Where Next

You now know

  • Logistic regression predicts P(y = 1 | x) = σ(w·x + b), a probability between 0 and 1.
  • It assumes the log-odds log(p / (1 − p)) are linear in x; ew is the odds ratio for a one-unit increase.
  • It is trained by maximum likelihood — equivalently, by minimizing the log-loss with gradient descent, whose gradient is (p − y)x.
  • A threshold turns probabilities into classes; the boundary w·x + b = 0 is linear; moving the threshold trades missed cases for false alarms.

Accuracy alone hid the missed patients in Section 7. Module 5 gives us the full toolkit: the confusion matrix, precision, recall, F1, ROC curves and AUC — plus the regression metrics.

Logistic Regression

Objectives 1. Classification 2. Types 3. Sigmoid 4. The Model 5. Odds & Logit 6. Max Likelihood 7. Threshold 8. Worked Example 9. Multiple Python Reading Exercises Recap