What it means for a computer to learn, the three ways it can learn, and how we know whether it learned anything useful.
Module 1 · Weeks 1–3 · Lecture notes by Dr. Abdulkarim Albanna
Foundations Concepts ~50 minTools for this course: Python in Google Colab (free, nothing to install) or a local IDE such as PyCharm, with numpy, pandas, matplotlib and scikit-learn.
Artificial Intelligence (AI) is the science of building systems that simulate and extend human intelligence. John McCarthy coined the term in 1956 and defined it as “the science and engineering of making intelligent machines, especially intelligent computer programs”.
Machine Learning (ML) is the part of AI that studies how computers can acquire knowledge or skills from data and improve their performance with experience, instead of being explicitly programmed with rules. An intelligent system that cannot learn is hard to call intelligent at all.
Deep Learning (DL) is the part of ML built on artificial neural networks with many hidden layers. It is the engine behind modern image, sound and text understanding.
A traditional program takes rules + data → answers. A machine-learning program takes data + answers → rules (a model). The model is then applied to new data it has never seen.
Learning, for people, is “the activity or process of gaining knowledge or skill by studying, practicing, being taught, or experiencing something” (Merriam-Webster). For machines we use Tom Mitchell's precise definition:
| Problem | Task T | Performance P | Experience E |
|---|---|---|---|
| Handwriting recognition | Recognize and classify handwritten words in images | % of words correctly classified | A database of labelled handwritten words |
| Robot driving | Drive on public highways using vision sensors | Average distance travelled before an error | Images and steering commands recorded from a human driver |
| Chess | Play chess | % of games won against opponents | Games played against itself |
| Spam filtering | Label e-mails as spam / not spam | % of e-mails correctly labelled | E-mails the user has marked as spam |
Whether the learner is a human or a machine, the process breaks into four parts:
“The last 10 years have been about building a world that is mobile-first. In the next 10 years we will shift to a world that is AI-first” — Sundar Pichai. You already use ML every day:
| Area | Examples |
|---|---|
| Computer vision | Face recognition, object detection, medical imaging (retina scans), self-driving cars |
| Speech & language | Speech-to-text, voice assistants, Google Translate, sentiment analysis, spam filtering |
| Recommendation | Netflix, Amazon and YouTube suggestions; Instagram feed ranking |
| Finance & business | Credit scoring, fraud detection, customer segmentation, algorithmic trading |
| Games & robotics | AlphaGo, Atari-playing agents, robot control |
| Generative AI | GANs and diffusion models that create realistic images; large language models |
In supervised learning every training example comes with the correct answer (a label). The algorithm learns a function that maps inputs to outputs, then predicts the label of new, unseen inputs. It is a predictive task: like a student learning from solved examples.
| Classification | Regression | |
|---|---|---|
| Output | A category (class) | A continuous number |
| Question | “Which class?” | “How much?” |
| Examples | Spam / not spam; cat / dog; admit / reject a student; disease / healthy | House price from size; exam score from study hours; temperature tomorrow |
| Algorithms in this course | Logistic regression, decision trees, Naïve Bayes, SVM, neural networks | Linear regression (Module 3) |
100 e-mails were flagged by hand: 25 spam and 75 not spam. These 100 labelled e-mails are the training set. A classifier learns which words and senders separate the two groups, and then labels each new e-mail on arrival. Because the output is one of two categories, this is binary classification.
In unsupervised learning the data has no labels. The algorithm looks for structure on its own — groups of similar items, unusual items, or items that occur together. It is a descriptive task: there is no correct answer to compare against.
Labelling is slow and expensive; unlabelled data is cheap. Semi-supervised learning uses a small labelled set plus a large unlabelled set. For example, to recommend web pages we would need users to mark pages they like — few will do it, but we have millions of unlabelled pages. The algorithm learns from the few labels and uses the structure of the unlabelled data to improve.
In reinforcement learning (RL) an agent learns by trial and error. It takes an action in an environment, receives a reward or punishment, and adjusts its behaviour to maximize the total reward over time — just as you train a dog with treats.
| Supervised | Unsupervised | Reinforcement | |
|---|---|---|---|
| Data | Labelled examples | Unlabelled examples | No dataset; interaction with an environment |
| Feedback | The correct answer | None | A delayed reward signal |
| Goal | Predict labels | Discover structure | Learn a policy that maximizes reward |
Every algorithm in this course builds a model — a summary of the data that lets us make predictions. Models fall into three families by how they describe the data:
| Family | Idea | Examples in this course |
|---|---|---|
| Logical | Divide the instance space with logical expressions (IF–THEN rules) | Decision trees (Module 6), rule sets |
| Geometric | Treat examples as points in space; use lines, planes and distances | Linear regression, SVM, k-nearest neighbours, k-means |
| Probabilistic | Treat features and targets as random variables; model the uncertainty with probabilities | Naïve Bayes (Module 7), logistic regression |
Probabilistic models are either discriminative (model P(Y|X) directly) or generative (model the joint P(Y, X) and use Bayes' rule):
A model is only useful if it works on new data. Two things can go wrong:
| Underfitting | Overfitting | |
|---|---|---|
| What happens | Model is too simple to capture the pattern | Model memorizes noise in the training data |
| Training accuracy | Low | Very high |
| Test accuracy | Low | Low |
| Error type | High bias | High variance |
| Typical fix | More features, a more flexible model, train longer | More data, a simpler model, regularization, pruning |
To estimate how a model will do on unseen data, we must test it on data it did not train on. The usual split:
When data is scarce, split it into k equal folds. Train on k−1 folds and validate on the remaining one; repeat k times so every fold is used once for validation, then average the k scores.
A dataset has 1,000 rows. A 50/25/25 split gives 500 training, 250 validation and 250 test rows. With 5-fold cross-validation on the 750 non-test rows, each fold has 750 / 5 = 150 rows: every model trains on 600 rows and validates on 150. If the five validation accuracies are 0.82, 0.85, 0.80, 0.84 and 0.84, the cross-validated accuracy is (0.82+0.85+0.80+0.84+0.84) / 5 = 0.83.
Module 5 covers the measures themselves (accuracy, precision, recall, F1, ROC/AUC, MAE, RMSE, R²).
Open a new notebook in Google Colab and run this. It follows the full workflow from this module: load labelled data, split it, train a classifier, and evaluate it on unseen data.
Change max_depth to 1 and then to 20. Which setting underfits (both accuracies low)? Which one makes the training accuracy perfect while the test accuracy stops improving?
Both books are free to read online from their authors: statlearning.com (T1) and probml.github.io (T2).
A bank wants a system that predicts whether a loan applicant will default. Write the task T, performance measure P and experience E.
T: classify each applicant as “will default” or “will repay”. P: % of applicants classified correctly (or better, recall on the defaulters, since missing a defaulter is costly). E: historical loan records of past applicants, each labelled with whether they defaulted.
Name the learning type for each: (a) grouping news articles by topic with no topic labels; (b) predicting tomorrow's temperature from past weather records; (c) teaching a robot to walk by rewarding forward progress; (d) detecting fraudulent card transactions from transactions already marked fraud / legitimate.
(a) Unsupervised (clustering). (b) Supervised — regression (numeric output). (c) Reinforcement learning. (d) Supervised — classification (fraud / legitimate).
(a) a student's final mark out of 100; (b) whether a student passes; (c) the number of visitors to a website tomorrow; (d) the blood type of a patient.
(a) Regression. (b) Classification (binary). (c) Regression (a count, treated as a number). (d) Classification (four classes: A, B, AB, O).
Model A: training accuracy 99%, test accuracy 71%. Model B: training accuracy 62%, test accuracy 60%. Which is overfitting and which is underfitting? Suggest one fix for each.
A overfits (high variance): the large train–test gap shows it memorized the training data. Fix: more data, a simpler model, or regularization/pruning. B underfits (high bias): both scores are low. Fix: a more flexible model or better features.
You run 4-fold cross-validation on 800 rows. How many rows are in each training run and each validation fold? The fold errors are 0.12, 0.10, 0.15 and 0.11. What is the cross-validated error?
Each fold has 800 / 4 = 200 rows, so each run trains on 600 and validates on 200. CV error = (0.12 + 0.10 + 0.15 + 0.11) / 4 = 0.48 / 4 = 0.12.
Classify each: (a) the slope of a regression line; (b) the depth of a decision tree; (c) the number of clusters k in k-means; (d) the weights of a neural network; (e) the learning rate.
Parameters (learned from data): (a) the slope and (d) the weights. Hyperparameters (chosen by us, tuned on the validation set): (b) tree depth, (c) k, and (e) the learning rate.
Every model is only as good as its data. Next, Module 2 covers how to clean and prepare a dataset before any learning happens.