Your assistant clusters colleges by tuition and admission rate, then asks whether K equals 2 or 4, and whether to bolt on five more features. Both questions turn on the same handful of clustering rules.
Why cluster at all. Clustering collapses many variables into a single group label. That label is a new feature you feed into a downstream model. It reveals hidden structure (customer segments, geographic zones) and reduces dimensionality without a response variable to guide it. That last point defines it as unsupervised.
KEY: Supervised learning has a target; unsupervised does not. Clustering finds groups, it does not predict a labeled outcome.
K-means, the algorithm. You pick K. The algorithm places K centroids, assigns each point to its nearest centroid, recomputes each centroid as the mean of its members, and repeats until assignments stop changing. It minimizes the total within-cluster sum of squared distances (WCSS).
Common mistakes
- Not standardizing features. Leaving tuition in dollars and admission rate as a fraction lets tuition dominate every distance. Center and scale first, or the clusters are meaningless.
- Listing converse pairs as two differences. "K-means fixes K" and "hierarchical does not fix K" is one difference. Graders collapse duplicates and converses; give genuinely separate points.
- Reading axis labels instead of explaining the tradeoff. For the elbow, higher K means tighter but less stable and less interpretable clusters; lower K means simpler but coarser. State the mechanism, not just "K=4 has lower WCSS."
Bottom line
- Clustering is unsupervised: it groups observations by similarity with no target variable.
- K-means needs you to fix the number of clusters K in advance; hierarchical does not.
- Standardize features first, or the largest-scale variable dominates the distance.
- Use an elbow plot (within-cluster sum of squares vs K) to choose K for K-means.
Exam shortcut
For any "similarities and differences" prompt, write your two differences on opposite axes (does K need pre-specifying? is output a tree or flat labels?) so graders cannot collapse them as converses. Before you interpret any cluster output, confirm the features were standardized; if the prompt gives raw dollars alongside rates, standardization is almost certainly the intended critique.
The full lesson (about 2,971 words, 20 min read) adds 2 worked examples, all 6 common mistakes, a self-check, free in the app.
Learning objectives
- 3c
Browse all free Exam PA lessons or jump into free Exam PA practice questions.