Your assistant proposes picking a boosted tree's learning rate by scanning the test-set curve for its lowest error. That instinct feels efficient, and it quietly leaks the test set into tuning. Knowing why separates a full-credit answer from a partial one.
What boosting does. Bagging and random forests build many deep trees in parallel and average them. Boosting instead builds trees one at a time. Every tree is shallow and weak. Each new tree focuses on the observations the current ensemble predicts poorly. You add trees slowly, and the ensemble improves in small steps.
The mechanism is sequential error correction. Start with a simple guess. Compute how wrong you are. Fit a small tree to those errors. Add a shrunken slice of it to the running prediction. Repeat. Because each tree targets the leftover mistakes, boosting drives down bias rather than variance.
Gradient boosting. In a gradient boosting machine, "errors" are made precise as the negative gradient of a loss function. For squared-error loss the gradient is just the residual, actual minus predicted.
Common mistakes
- Tuning on the test set. Choosing the learning rate or tree count by minimizing test error leaks test data into model selection. Use cross-validation on the training data and reserve the test set for one final estimate.
- Reading a high learning rate as merely faster. A learning rate of 0.30 trains quickly but overfits; the falling train error with rising test error is memorization, not efficiency.
- Fixing tree count while changing the rate. Dropping from 0.10 to 0.05 but keeping 500 trees underfits. Lower rates need proportionally more trees.
Bottom line
- Boosting fits trees sequentially; each new tree corrects the residual errors of the ensemble so far.
- Gradient boosting fits each tree to the negative gradient of the loss, scaled by the learning rate (shrinkage).
- Small learning rate plus many trees generalizes better; large learning rate overfits fast.
- Learning rate and number of trees trade off inversely: halve the rate, roughly double the trees.
Exam shortcut
When a prompt mentions choosing hyperparameters from the test set, immediately write "data leakage" and recommend cross-validation on the training data; that phrase is the graders' target. To reason about learning rate and tree count, remember they move inversely: smaller rate demands more trees, and halving the rate roughly doubles the trees.
The full lesson (about 1,862 words, 12 min read) adds 2 worked examples, all 6 common mistakes, a self-check, free in the app.
Learning objectives
- 5c
Browse all free Exam PA lessons or jump into free Exam PA practice questions.