A hospital wants a model to flag patients likely to be readmitted within 30 days. Before you write a line of code, four gates decide whether the project is worth starting at all.
Why gate a project at all. A predictive model earns its cost only when a better forecast changes a decision that matters. Screening upfront stops you from building an accurate model no one can use. Weigh four factors in order.
Available data. Confirm a target variable exists with recorded historical outcomes. Then confirm predictors are relevant, reasonably clean, and voluminous enough to learn a signal.
DECISION: If no historical outcome variable exists, you cannot fit a supervised model yet. The honest next step is data collection, not modeling.
TRAP: "We have lots of data" is not enough. Volume without a labeled target, or predictors unrelated to the outcome, still blocks a predictive model.
Available technology. A model that lives only in a notebook creates no value. Ask whether the firm can score new records, deploy predictions into the workflow, and refresh the model...
Common mistakes
- Equating data volume with readiness. A large table with no target label still cannot train a supervised model.
- Skipping the deployment question. A model with no path to production delivers zero business impact regardless of accuracy.
- Chasing accuracy over impact. Improving error from 12% to 10% is worthless if that gain changes no decision.
Bottom line
- Screen every proposed model against four factors before committing: data, technology, business impact, and implementation.
- Data means a usable target variable plus relevant, clean predictors at adequate volume.
- No historical outcome label means no supervised model yet; collect data first.
- Technology means the infrastructure to build, deploy, and refresh the model in production.
Exam shortcut
Walk the four gates in order (data, technology, business impact, implementation) and state which pass and which fail; a single failed gate can block the project. For any "should we build this?" prompt, first confirm a historical outcome label exists, because its absence forces "collect data" as the answer. When impact is questioned, tie it to a specific decision and its dollar size, not to model accuracy alone.
The full lesson (about 1,283 words, 9 min read) adds 2 worked examples, all 6 common mistakes, a self-check, free in the app.
Learning objectives
- 1e
Browse all free Exam PA lessons or jump into free Exam PA practice questions.