A property model with 200 levels of building type and three square-footage columns that nearly duplicate each other will overfit and produce unstable coefficients. Regularization is the tool that tames both problems, and the exam tests whether you can tune it.
Why penalize at all. Ordinary least squares minimizes squared error and nothing else. With many predictors, many-level categoricals, or correlated columns, ordinary least squares (OLS) chases noise and coefficient estimates swing wildly. Regularization adds a penalty term that grows as coefficients grow. The model must now balance fitting the data against keeping coefficients small.
KEY: The penalty deliberately introduces bias. You accept slightly biased coefficients in exchange for far lower variance, which usually lowers test-set error. This is the bias-variance trade-off made tunable.
Regularization is also the natural cure for multicollinearity. When predictors are highly correlated, OLS cannot separate their effects, so standard errors explode and signs flip. The variance-inflation factor measures this.
Common mistakes
- Claiming ridge sets coefficients to zero. Ridge shrinks but never zeroes. Only lasso and elastic net produce exact zeros. Saying "ridge drops variables" is a classic point-loser.
- Confusing lambda with alpha. Lambda is penalty strength (0 gives OLS, large gives intercept-only). Alpha is the L1-L2 mix (0 = ridge, 1 = lasso). Graders flagged candidates who called a "large lambda" equal to 1, confusing it with the alpha mixing weight.
- Misreading the lambda extremes. At lambda = 0 the model is plain OLS. As lambda grows very large, all slopes shrink toward zero and the fit approaches the mean (intercept only). Full credit required both extremes stated correctly.
Bottom line
- Regularization adds a penalty on coefficient size to the loss, trading a little bias for a large drop in variance.
- Ridge uses an L2 penalty (sum of squared coefficients); it shrinks toward zero but never exactly to zero, so all predictors stay in.
- Lasso uses an L1 penalty (sum of absolute coefficients); it can force coefficients to exactly zero, performing variable selection.
- Elastic net blends L1 and L2 through the mixing parameter alpha, keeping lasso selection while handling correlated groups like ridge.
Exam shortcut
Map the letters fast: L1 is lasso and it zeros coefficients (one absolute value has a corner at zero); L2 is ridge and it only shrinks. When a task asks about lambda extremes, always write both ends: lambda = 0 is OLS, huge lambda is intercept-only. For any "recommend a method for interpretability" prompt, pick backward selection, lasso, or elastic net (never ridge) and justify by reduced coefficient count.
The full lesson (about 2,243 words, 15 min read) adds 2 worked examples, all 6 common mistakes, a self-check, free in the app.
Learning objectives
- 4d
Browse all free Exam PA lessons or jump into free Exam PA practice questions.