A single pruned tree is easy to read but jumpy: shift a few training rows and the top split flips. Bagging and random forests trade that instability for accuracy by averaging many trees.
Why average trees at all. A deep tree has low bias but high variance. Averaging many independent predictions cuts variance without raising bias. That is the whole engine behind both methods.
Bagging. Draw B bootstrap samples (sample n rows with replacement). Fit one unpruned tree per sample. To predict, average the B regression outputs or take the majority class.
Push high and the second term vanishes, but the first term stays. That floor is bagging's weakness: bootstrap trees are highly correlated because one or two strong predictors dominate the top split in every tree.
KEY: Random forests attack the correlation term , not the count . At each split they consider a random subset of mtry predictors, so strong predictors cannot own every tree. Lower , lower variance.
Common mistakes
- Averaging only one tree. October 2023 graders flagged candidates who used a single tree when the task asked for a bagged prediction. A bagged forecast averages all trees; using one is not bagging.
- Forgetting the absolute value. In MAE work, |prediction − actual| must be taken before comparing. Skipping it (leaving −$500 instead of $500) was a named grader complaint.
- Thinking more trees removes all variance. Raising only kills the term. The floor stays; only lowering mtry-driven correlation removes it.
Bottom line
- Bagging is bootstrap aggregating: fit a tree on each of B bootstrap resamples, then average predictions (regression) or majority-vote (classification).
- Bagging reduces variance, not bias, so grow deep unpruned trees as the base learners.
- Random forests improve bagging by sampling only mtry predictors at each split, which decorrelates the trees and lowers variance further.
- Default mtry is about sqrt(p) for classification and p/3 for regression; tune it by cross-validation.
Exam shortcut
For any bagging arithmetic task, average every tree first, then take the absolute value before comparing errors; the two most common lost points are using one tree and dropping the absolute value. When recommending mtry (or any random forest parameter), read the tuning plot, name the value at the error minimum, and add one sentence on why it beats its neighbors.
The full lesson (about 1,654 words, 11 min read) adds 2 worked examples, all 6 common mistakes, a self-check, free in the app.
Learning objectives
- 5b
Browse all free Exam PA lessons or jump into free Exam PA practice questions.