Sample Questions
The validation set approach randomly divides the dataset into two parts: a training set (used to fit the model) and a validation set (used to evaluate the model's performance). The model is trained on the training portion and its error is estimated on the held-out validation portion.
The Bayes classifier assigns each observation to the most probable class given the observed features: . This is the theoretically optimal classifier that minimizes the overall misclassification rate.
Bias measures the systematic error introduced by the modeling assumptions. It is the difference between the expected prediction (averaged over many training sets) and the true function value . A high-bias model makes strong assumptions that may not match the true relationship (underfitting).