Course outline
Chapter 1: Supervised Learning
Linear Regression
Model, loss function, the normal equation and gradient descent — the foundation for every model that follows.
Linear regression predicts a continuous target from a feature vector . It is simple, fast and easy to interpret, and almost every idea in this course (loss functions, optimisation, regularisation) appears here first.
The model
Given training examples with , we assume
where we prepend to every example so the intercept fits into the dot product.
Stacking all examples as rows gives the design matrix and the prediction vector .
Loss function
We measure how wrong the model is with the mean squared error (MSE):
The factor is only there to cancel the 2 that appears when we differentiate.
Solving for the parameters
The normal equation
is a convex quadratic, so its minimum is where the gradient vanishes:
This is exact but costs to invert , which becomes slow when is large.
Gradient descent
Instead of solving in one step, we repeatedly move against the gradient with a learning rate :
Implementation in Python
import numpy as np
def fit_normal_equation(X: np.ndarray, y: np.ndarray) -> np.ndarray:
"""Closed-form solution. X must already include a column of ones."""
return np.linalg.solve(X.T @ X, X.T @ y)
def fit_gradient_descent(
X: np.ndarray, y: np.ndarray, lr: float = 0.1, epochs: int = 1_000
) -> np.ndarray:
n, d = X.shape
theta = np.zeros(d)
for _ in range(epochs):
grad = X.T @ (X @ theta - y) / n
theta -= lr * grad
return theta
rng = np.random.default_rng(0)
x = rng.uniform(-1, 1, size=200)
y = 3.0 + 2.0 * x + rng.normal(scale=0.1, size=200)
X = np.column_stack([np.ones_like(x), x])
print(fit_normal_equation(X, y)) # ≈ [3.0, 2.0]
print(fit_gradient_descent(X, y)) # ≈ [3.0, 2.0]
Notice that we use np.linalg.solve instead of computing an explicit inverse — it is faster and numerically more stable.
Evaluating the model
The coefficient of determination tells us what fraction of the variance in the model explains:
| Metric | Formula | Interpretation |
|---|---|---|
| MSE | Average squared error, in squared units of | |
| RMSE | Same units as | |
| see above | 1 is perfect, 0 is no better than predicting |
Summary
- Linear regression models and minimises the MSE.
- The normal equation gives an exact solution; gradient descent scales to large .
- Always scale features and evaluate on held-out data.
Exercise. Show that is invertible if and only if the columns of are linearly independent. What happens to the normal equation when two features are identical?