A block-level energy model has data from five years ago and a rare-fire target that appears in 2% of rows. Before you fit anything, three design choices decide whether the model is trustworthy: how far back you reach, how you sample, and how finely you slice your factors.
Time frame. Every historical dataset forces a trade. Reach further back and you gain rows, which lowers variance and stabilizes estimates. Reach too far and you pull in periods that no longer represent today: a repriced product, a new underwriting rule, a regulatory shift. Those older values bias the model toward conditions that will not recur.
KEY: More history is not automatically better. Include an older period only if the data-generating process was materially the same. Otherwise the extra rows introduce bias that outweighs the variance reduction.
The judgment is contextual. A slow-moving mortality process tolerates a long window. A fast-moving pricing or fraud process does not. State the trade explicitly on the exam: older data helps sample size but may not be representative.
Common mistakes
- Assuming more history is always better. Older rows help sample size but bias the model if the process changed. State the representativeness trade, not just the volume gain.
- Oversampling before the split. Duplicated minority rows leak across the train/test boundary and inflate test performance. Always sample within training only.
- Calling oversampling a proportional-representation method. It rebalances a target; it does not preserve population proportions. That is stratified sampling's job.
Bottom line
- Time frame trades data volume against representativeness: older rows add sample size but may reflect conditions that no longer hold.
- Random sampling gives every row an equal chance; stratified sampling forces proportional representation within defined subgroups.
- Systematic sampling takes every k-th row; it fails if the ordering has a hidden cycle.
- Oversampling duplicates the rare class; undersampling drops majority rows. Both rebalance an imbalanced target.
Exam shortcut
For any time-frame question, write one sentence both ways: older data raises sample size but may not be representative of current conditions. For imbalance, name the rare-class rate, note that accuracy misleads, recommend oversampling or undersampling, and add the split-first rule to preempt the leakage trap. For granularity critiques, ask a single question: does this variable vary across the units I am predicting?
The full lesson (about 1,817 words, 12 min read) adds 2 worked examples, all 6 common mistakes, a self-check, free in the app.
Learning objectives
- 2c
Browse all free Exam PA lessons or jump into free Exam PA practice questions.