Exam PA · Tree-Based Models · Free Lesson

Construct, prune, and validate regression and classification trees.

Free SOA Exam PA (Predictive Analytics) lesson in Tree-Based Models. 15 min read, ~2,247 words.

Your assistant hands you an overgrown tree that scores perfectly on training data and poorly on the holdout. The fix is not a new model. It is knowing how splits are chosen, how the complexity parameter prunes, and how a confusion matrix tells you it was overfit.

How a split is chosen. A tree searches every predictor and every candidate cutpoint, then keeps the split that most improves node purity. For classification, purity is measured two ways.

Entropy measures disorder. A node with all one class has entropy zero. A perfect 50/50 split has entropy one (in bits).

H=−∑kpklog⁡2pkH = -\sum_{k} p_k \log_2 p_k

Gini impurity measures the chance of misclassifying a random observation labeled by the node's class distribution.

G=1−∑kpk2G = 1 - \sum_{k} p_k^{2}

Both are zero at a pure node and peak at an even split. rpart uses Gini by default. The split with the largest weighted drop in impurity (the information gain) wins.

Read the full lesson, free →
Worked examples and practice. Free with a free account, no card.

Common mistakes

Bottom line

Exam shortcut

Gini is faster to compute by hand than entropy: 1−∑pk21 - \sum p_k^2, no logs. For information gain, always subtract the observation-weighted child impurity from the parent, never the simple average. When asked to fight overfitting, reach for one lever and state its direction: raise cp, raise minbucket, or lower maxdepth.

The full lesson (about 2,247 words, 15 min read) adds 2 worked examples, all 6 common mistakes, a self-check, free in the app.

Learning objectives

Browse all free Exam PA lessons or jump into free Exam PA practice questions.