Skip to content
Course outline

Chapter 2: Classification

Logistic Regression

From linear scores to probabilities with the sigmoid function and the cross-entropy loss.

The sigmoid function

To turn a real-valued score into a probability we use

σ(z)=11+e−z,P(y=1∣x)=σ(θ⊤x).\sigma(z) = \frac{1}{1 + e^{-z}}, \qquad P(y = 1 \mid \mathbf{x}) = \sigma(\boldsymbol\theta^\top \mathbf{x}).

Cross-entropy loss

J(θ)=−1n∑i=1n[y(i)log⁡y^(i)+(1−y(i))log⁡(1−y^(i))]J(\boldsymbol\theta) = -\frac{1}{n}\sum_{i=1}^{n} \Big[ y^{(i)} \log \hat{y}^{(i)} + (1 - y^{(i)}) \log (1 - \hat{y}^{(i)}) \Big]

Its gradient has the same form as in linear regression: ∇J=1nX⊤(y^−y)\nabla J = \frac{1}{n} X^\top(\hat{\mathbf{y}} - \mathbf{y}).

In scikit-learn

from sklearn.linear_model import LogisticRegression

model = LogisticRegression(max_iter=1_000)
model.fit(X_train, y_train)
print(model.score(X_test, y_test))