An underwriter hands you claim counts, a binary fraud flag, and a column of dollar severities, then asks for one modeling approach. Each column demands a different generalized linear model (GLM), and naming the right family plus link earns the point.
Match the family to the response. A GLM has three pieces: a random component (the response distribution), a systematic component (the linear predictor ), and a link function connecting the mean to the linear predictor. The single most tested skill here is choosing the distribution and link that fit the data-generating process.
KEY: Two questions pick the model. What does the response look like (bounded 0/1, integer count, positive skewed)? And is the effect additive or multiplicative? Bounded and skewed responses rule out plain ordinary least squares (OLS); multiplicative effects call for a log or logit link.
Why not just use OLS everywhere? Ordinary least squares assumes a normally distributed response with constant variance and an unbounded range. Claim counts are non-negative integers. A fraud flag is bounded in .
Common mistakes
- Exponentiating a logit coefficient into a probability. For binary GLMs, is an odds ratio, not a change in probability. A coefficient of gives an odds ratio of , which is not a 65% rise in the probability of the event.
- Using plain Poisson on overdispersed counts. If the sample variance far exceeds the mean, plain Poisson understates standard errors. The tell is a dispersion estimate well above one; switch to negative binomial or quasi-Poisson.
- Reading raw residual plots for a GLM. Raw residuals have mean-dependent variance, so a fan shape is expected and uninformative. Use deviance residuals, which R plots by default.
Bottom line
- Pick the family from the response type: binary is binomial, counts are Poisson, non-negative skewed amounts are Gamma.
- The log link makes coefficients multiplicative; exp(beta) is the factor by which a one-unit rise scales the expected response.
- The logit link makes exp(beta) an odds ratio, not a probability change.
- A normal regression on log(Y) models the mean of the log; a GLM log link models the log of the mean. They differ.
Exam shortcut
Read the response column before anything else: 0/1 means binomial with logit, integer counts mean Poisson with log, positive skewed dollars mean Gamma with log. Naming the pair fast banks easy points. For any log-link coefficient, immediately write and read it as a multiplier; for a logit coefficient, say "odds ratio" out loud so you never call it a probability.
The full lesson (about 2,204 words, 15 min read) adds 2 worked examples, all 6 common mistakes, a self-check, free in the app.
Learning objectives
- 4a
Browse all free Exam PA lessons or jump into free Exam PA practice questions.