A dataset arrives with raw claim amounts, a birth date, and a five digit ZIP code. None of those columns is ready to model. Turning them into a log severity, an age band, and a region indicator is feature engineering, and the exam rewards the candidate who recommends the right transform and justifies it.
Variables versus features. A variable is what the source system recorded: a timestamp, a dollar amount, a text label. A feature is the input you actually give the model, engineered so the signal is easy to extract. Every feature starts as one or more variables, but not every variable is a good feature. Birth date is a variable; age in years, or an age band, is the feature.
Why generate features. Raw data rarely matches a model's assumptions. A linear model wants roughly linear, comparably scaled inputs. A tree wants informative splits. Feature generation closes that gap.
Single function transforms. Apply one function to one variable. The log is the workhorse. Right skewed, strictly positive quantities like claim severity, income, or population compress toward symmetry under a...
Common mistakes
- Scaling a skewed variable and calling it fixed. Min max scaling a cost with a value of $5,000 outlier still leaves skew and the outlier; only a transform like log reshapes the distribution.
- Summing a level or averaging a flow. Averaging monthly kWh or summing daily temperatures (30 days at 84 giving 2,520) produces a meaningless feature. Average levels, sum flows.
- Feeding a raw address as a factor. Any address not in training data cannot be predicted. Extract ZIP, borough, or region instead of using the full string.
Bottom line
- A variable is a raw recorded column; a feature is a variable reshaped to help a model learn.
- Feature generation adds value when raw fields are skewed, high cardinality, non linear, or not on a comparable scale.
- Single function transforms (log, square root) compress skew and linearize multiplicative effects.
- One variable into many: binarization (0/1 flag) and bucketing (binning a numeric into ordered bands).
Exam shortcut
To pick an aggregation, ask "level or flow?" Average levels and rates; sum counts and totals, and justify with the variable's nature for full credit. When a prompt says standardize or scale, immediately check whether skew or outliers are present; if they are, state that scaling does not fix them and name log as the transform that does.
The full lesson (about 2,024 words, 13 min read) adds 2 worked examples, all 6 common mistakes, a self-check, free in the app.
Learning objectives
- 3a
Browse all free Exam PA lessons or jump into free Exam PA practice questions.