Free SOA Exam PA (Predictive Analytics) Formula Sheet (2026)

Every Exam PA formula you need on the test, grouped by topic and rendered with full math notation. 53 formulas across 5 topics, calibrated to the 2026 syllabus. Free forever, no signup required.

53 formulas 5 topics 2026 syllabus Free forever
Print-ready PDF: 1080x1350 portrait, math pre-rendered, fonts embedded. Download once, study anywhere.
Download PDF →

All Exam PA Formulas

Predictive Analytics Problem Definition 7 items
Profit per policy
Profit=P−L−E\text{Profit} = P - L - E, P = premium (price) per policy, L = expected loss (claim severity times frequency), E = expense or acquisition cost per policy
Value as price minus cost
V=P−CV = P - C, V = value of the item, P = price (e.g. sale price per building), C = cost (e.g. construction cost per building)
Base rate of a binary target
p=neventNp = \dfrac{n_{event}}{N}, p = base rate (positive-class proportion), neventn_{event} = number of positive-class records, N = total records
Majority-class (no-information) accuracy
accmaj=1−pacc_{maj} = 1 - p, accmajacc_{maj} = accuracy of always predicting the majority class, p = base rate of the rare (positive) class
Business impact of a predictive model
Impact=V×f\text{Impact} = V \times f, V = dollar value of the decision the prediction improves, f = how often that decision is made
Additive error data-generating model
y=f(x)+εy = f(x) + \varepsilon, y = observed response, f(x) = true signal function, ε\varepsilon = random noise with variance Var⁡(ε)\operatorname{Var}(\varepsilon)
Bias-variance decomposition of expected test MSE
E[(y0−f^(x0))2]=Var⁡(f^(x0))+[Bias⁡(f^(x0))]2+Var⁡(ε)E\left[(y_0 - \hat f(x_0))^2\right] = \operatorname{Var}(\hat f(x_0)) + \left[\operatorname{Bias}(\hat f(x_0))\right]^2 + \operatorname{Var}(\varepsilon), f^\hat f = fitted model, x0x_0 = test point, y0y_0 = true response, Var⁡(ε)\operatorname{Var}(\varepsilon) = irreducible error
Data Exploration and Visualization 8 items
Variance of a factor-level estimate
Var⁡(μ^level)∝1nlevel\operatorname{Var}(\hat\mu_{\text{level}}) \propto \frac{1}{n_{\text{level}}}, μ^level\hat\mu_{\text{level}} = estimated mean for that level, nleveln_{\text{level}} = number of rows in that level
Minority rows needed to oversample to a target share
f=pN1−pf = \frac{pN}{1-p} from p=fN+fp = \frac{f}{N+f}, p = target minority proportion, N = majority-class row count, f = minority rows after oversampling
Boxplot outlier fences
lower=Q1−1.5×IQR, upper=Q3+1.5×IQR\text{lower} = Q1 - 1.5\times\text{IQR},\ \text{upper} = Q3 + 1.5\times\text{IQR}, Q1 = first quartile, Q3 = third quartile, IQR = interquartile range
Interquartile range
IQR=Q3−Q1\text{IQR} = Q3 - Q1, Q1 = 25th percentile (first quartile), Q3 = 75th percentile (third quartile)
Shared variance between two predictors
r2=rxy2r^2 = r_{xy}^2, rxyr_{xy} = Pearson correlation between the two variables; r2r^2 = proportion of one variable's variation explained by the other
Sample variance
s2=1n−1∑i=1n(xi−xˉ)2s^{2} = \frac{1}{n-1}\sum_{i=1}^{n}(x_i - \bar{x})^{2}, xix_i = observation, xˉ\bar{x} = sample mean, n = sample size
Pearson correlation coefficient
rxy=∑(xi−xˉ)(yi−yˉ)∑(xi−xˉ)2∑(yi−yˉ)2r_{xy} = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2}\sqrt{\sum (y_i - \bar{y})^2}}, x,y = paired values, xˉ,yˉ\bar{x},\bar{y} = means, i = observation index
Sample standard deviation
s=s2=1n−1∑i=1n(xi−xˉ)2s = \sqrt{s^{2}} = \sqrt{\frac{1}{n-1}\sum_{i=1}^{n}(x_i - \bar{x})^{2}}, s2s^2 = sample variance, xix_i = observation, xˉ\bar{x} = mean, n = sample size
Data Transformations and Unsupervised Learning Techniques 12 items
Standardized feature (z-score for clustering)
z=x−μσz = \dfrac{x - \mu}{\sigma}, x = raw feature value, μ = feature mean, σ = feature standard deviation
Complete linkage between-cluster distance
d(A,B)=max⁡a∈A, b∈B∥a−b∥d(A,B) = \max_{a \in A,\, b \in B} \lVert a - b \rVert, A, B = clusters, a, b = member points, distance = farthest pair
Proportion of variance explained by a principal component
PVEm=Var(zm)∑jVar(zj)\text{PVE}_m = \frac{\text{Var}(z_m)}{\sum_j \text{Var}(z_j)}, PVEm = proportion for PC m, Var(zm) = variance of PC m, denominator = total variance across all PCs
Variance of a principal component from its standard deviation
Var(zm)=sm2\text{Var}(z_m) = s_m^2, Var(zm) = variance captured by PC m, sm = reported standard deviation of PC m; total variance = number of standardized variables
Within-cluster sum of squares (K-means objective)
WCSS=∑k=1K∑x∈Ck∥x−μk∥2\text{WCSS} = \sum_{k=1}^{K} \sum_{x \in C_k} \lVert x - \mu_k \rVert^{2}, K = number of clusters, CkC_k = cluster k, x = observation, μkμ_k = mean vector of cluster k
Single linkage between-cluster distance
d(A,B)=min⁡a∈A, b∈B∥a−b∥d(A,B) = \min_{a \in A,\, b \in B} \lVert a - b \rVert, A, B = clusters, a, b = member points, distance = closest pair
First principal component score
z1=ϕ11x1+ϕ21x2+ϕ31x3z_1 = \phi_{11}x_1 + \phi_{21}x_2 + \phi_{31}x_3, z1 = PC1 score, φj1 = loading of variable j on PC1, xj = standardized feature j
Loading vector normalization constraint
∑jϕjm2=1\sum_{j} \phi_{jm}^2 = 1, φjm = loading of variable j on PC m; squared loadings of a component sum to one (unit-length loading vector)
Number of dummy columns for a categorical variable
d=k−1d = k - 1, k = number of category levels, d = number of dummy (indicator) columns needed to avoid redundancy in a GLM
Log transform for a skewed non-negative variable
x′=ln⁡(x+1)x' = \ln(x + 1), x = raw strictly non-negative value, x' = transformed value; the +1 lets x = 0 be handled
Min-max scaling
xmm=x−xmin⁡xmax⁡−xmin⁡x_{\text{mm}} = \dfrac{x - x_{\min}}{x_{\max} - x_{\min}}, x = raw value, xmin⁡x_{\min} = minimum, xmax⁡x_{\max} = maximum, xmmx_{\text{mm}} = scaled value in [0,1]
Z-score standardization
z=x−xˉsz = \dfrac{x - \bar{x}}{s}, x = raw value, xˉ\bar{x} = sample mean, s = sample standard deviation, z = standardized value
Generalized Linear Models 15 items
Poisson log-link model with a log-exposure offset
log⁡(E[count])=log⁡(exposure)+β0+β1x1+⋯\log(E[\text{count}]) = \log(\text{exposure}) + \beta_0 + \beta_1 x_1 + \cdots, exposure = units of exposure (offset, coef fixed at 1), β0\beta_0 = intercept, βj\beta_j = slope, xjx_j = predictor
Weighted least squares objective
min⁡β∑iwi(yi−y^i)2\min_{\beta} \sum_{i} w_i (y_i - \hat{y}_i)^2, wiw_i = weight of row i, yiy_i = observed response, y^i\hat{y}_i = fitted value, β\beta = coefficients
Group-size weighted mean of grouped rates
yˉw=∑iniyi∑ini\bar{y}_w = \dfrac{\sum_i n_i y_i}{\sum_i n_i}, nin_i = group size (weight) of row i, yiy_i = group-average rate for row i
Expected count as exposure times a fitted rate
E[count]=exposure×eβ0+β1x1E[\text{count}] = \text{exposure} \times e^{\beta_0 + \beta_1 x_1}, exposure = units of exposure, eβ0+β1x1e^{\beta_0 + \beta_1 x_1} = rate per unit exposure, β0\beta_0 = intercept, β1\beta_1 = slope, x1x_1 = predictor
GLM log link function
log⁡(μ)=β0+β1x1+…\log(\mu) = \beta_0 + \beta_1 x_1 + \dots, so eβ1e^{\beta_1} is the factor by which a one-unit rise in x1x_1 scales μ\mu; μ\mu = expected response, β\beta = coefficients
Pearson residual for a GLM
ri=yi−μ^iV(μ^i)r_i = \frac{y_i - \hat{\mu}_i}{\sqrt{V(\hat{\mu}_i)}}, yiy_i = observed, μ^i\hat{\mu}_i = fitted mean, V(μ^i)V(\hat{\mu}_i) = variance function at the fitted mean
Lognormal mean back-transformation (smearing correction)
E^[Y]=eμ^log⁡ eσ2/2\hat{E}[Y] = e^{\hat{\mu}_{\log}}\, e^{\sigma^2/2}, μ^log⁡\hat{\mu}_{\log} = fitted value on log scale, σ2\sigma^2 = residual variance on log scale
Estimated dispersion parameter
ϕ^=1n−p∑i=1n(yi−μ^i)2V(μ^i)\hat{\phi} = \frac{1}{n-p}\sum_{i=1}^{n}\frac{(y_i - \hat{\mu}_i)^2}{V(\hat{\mu}_i)}, n = observations, p = parameters, yiy_i = observed, μ^i\hat{\mu}_i = fitted mean, VV = variance function
Elastic net penalized regression objective
min⁡β12n∑i(yi−β0−∑jβjxij)2+λ[1−α2∑jβj2+α∑j∣βj∣]\min_{\beta} \frac{1}{2n}\sum_i (y_i - \beta_0 - \sum_j \beta_j x_{ij})^2 + \lambda[\frac{1-\alpha}{2}\sum_j \beta_j^2 + \alpha\sum_j |\beta_j|], λ = penalty strength, α = L1/L2 mix, β = coefficients
Variance-inflation factor for a predictor
VIFj=11−Rj2\text{VIF}_j = \frac{1}{1 - R_j^2}, Rj2R_j^2 = R-squared from regressing predictor j on all other predictors
One-standard-error threshold for lambda.1se selection
threshold=MSEmin⁡+SEmin⁡\text{threshold} = \text{MSE}_{\min} + \text{SE}_{\min}, MSEminMSE_{min} = minimum CV mean squared error, SEminSE_{min} = its standard error; lambda.1se = largest λ with CV MSE ≤ threshold
Log-link multiplicative effect of a coefficient
μnew=μ×eβ\mu_{new} = \mu \times e^{\beta} per unit, μ = expected response, β = coefficient on the link scale; eβ−1e^{\beta}-1 is the percent change
Logit-link odds ratio from a coefficient
OR=eβOR = e^{\beta}, OR = multiplicative change in the odds of the event per unit increase, β = logit-link coefficient
Interaction-adjusted slope of a predictor at a level
slope of x1 at level ℓ=β1+βint,ℓ\text{slope of } x_1 \text{ at level } \ell = \beta_1 + \beta_{\text{int},\ell}, β1 = main coefficient, βintβ_{int},ℓ = interaction coefficient (zero for baseline level)
GLM linear predictor with link function
g(μ)=η=β0+β1x1+…g(\mu) = \eta = \beta_0 + \beta_1 x_1 + \dots, g = link function, μ = expected response, η = linear predictor, β = coefficients, x = predictors
Tree-Based Models 11 items
Default mtry for random forest classification
mtry=pm_{try} = \sqrt{p}, p = total number of predictors, mtrym_{try} = predictors sampled at each split
Out-of-bag probability a row is left out of a bootstrap sample
P=(1−1/n)n→e−1≈0.368P = (1-1/n)^{n} \to e^{-1} \approx 0.368, n = number of training rows, P = chance a given row is omitted from one bootstrap (~37%)
Default mtry for random forest regression
mtry=p/3m_{try} = p/3, p = total number of predictors, mtrym_{try} = predictors sampled at each split
Gradient boosting additive update
Fm(x)=Fm−1(x)+λ hm(x)F_m(x) = F_{m-1}(x) + \lambda\, h_m(x), F = ensemble prediction, m = boosting round, λ = learning rate (shrinkage), hmh_m = tree fit to current pseudo-residuals
k-fold cross-validation error
CV error=1k∑i=1kError(foldi)\text{CV error} = \frac{1}{k}\sum_{i=1}^{k} \text{Error}(\text{fold}_i), k = number of folds, foldifold_i = i-th held-out validation fold
Variance of an averaged tree ensemble
Var⁡=ρσ2+1−ρBσ2\operatorname{Var} = \rho\sigma^{2} + \frac{1-\rho}{B}\sigma^{2}, ρ = pairwise tree correlation, σ² = single-tree prediction variance, B = number of trees
Learning-rate to tree-count inverse tradeoff
nnew≈nold×λoldλnewn_{new} \approx n_{old} \times \frac{\lambda_{old}}{\lambda_{new}}, n = number of trees, λ = learning rate; halving λ roughly doubles the trees needed
Cost-complexity pruning criterion
Rcp(T)=R(T)+cp⋅∣T∣R_{cp}(T) = R(T) + cp \cdot |T|, R(T) = tree total error, cp = complexity parameter, |T| = number of terminal nodes (leaves)
Gini impurity of a node
G=1−∑kpk2G = 1 - \sum_{k} p_k^{2}, G = Gini impurity, pkp_k = proportion of class k in the node; G = 0 at a pure node (rpart default)
Entropy of a node
H=−∑kpklog⁡2pkH = -\sum_{k} p_k \log_2 p_k, H = entropy in bits, pkp_k = proportion of class k in the node; H = 0 at a pure node
Root mean square error
RMSE=1n∑i(yi−y^i)2\text{RMSE} = \sqrt{\frac{1}{n}\sum_i (y_i - \hat{y}_i)^2}, yiy_i = actual value, y^i\hat{y}_i = predicted value, n = number of observations
Take the free Exam PA diagnostic quiz →
Instant readiness score in about 6 minutes. No signup to start.

Frequently Asked Questions

Is the Exam PA formula sheet free?
Yes. The full Exam PA formula sheet is free, with no signup, no email, and no credit card required. 53 formulas across 5 topics, all rendered with the same KaTeX math notation used in the FreeFellow study app.
Can I download the Exam PA formula sheet as a printable PDF?
Yes. A 1080x1350 portrait PDF (Instagram and LinkedIn carousel native size, also great for tablet study) is linked at the top of this page. The PDF is fully self-contained: math is pre-rendered, fonts are embedded, no internet connection needed once downloaded.
What's covered on the Exam PA formula sheet?
Every formula is grouped by official syllabus topic, with the formula in math notation plus a one-line note on when to use it (or a watch-out from CAIA, CFA, or other prep-provider commentary). Coverage is calibrated to the 2026 syllabus and refreshed when the corpus changes.
What is FreeFellow's relationship with SOA?
No. FreeFellow is not affiliated with the SOA or any examination body. This is an independent study aid covering the published syllabus.
What else is free at FreeFellow for Exam PA candidates?
The full original question bank is free with an account, subject to usage limits. Worked solutions, written lessons, mixed practice, and your readiness score stay free. The formula sheet is free too. Fellow is $39 per month or $79 per quarter, per exam family (USD). Fellow Plus is $49 per month, $99 per quarter, or $199 per year, per exam family (USD). Every annual plan is Fellow Plus. Fellow adds timed mock exams, spaced-repetition flashcards, performance analytics, and a personalized study plan. AI grading: 5 attempts a day on Fellow; Fellow Plus removes that allowance, subject to grading rate and usage limits.

About FreeFellow

Jeffrey Ting, founder of FreeFellow
Jeffrey Ting
FSA, CFA · Founder

FreeFellow was built by Jeffrey Ting, a credentialed actuary and CFA charterholder who has passed thirteen of the hardest exams in finance, all on his first attempt. He paid four-figure prep fees along the way.

So he started writing his own questions, then lessons, then mock exams, until it grew into a full prep platform covering 40 exams with more than 45,000 original practice questions.

01
Cost shouldn't decide who gets in.

A CFA charter, a CPA license, an actuarial credential. Each opens a real career, but prep fees add to the cost of getting there. FreeFellow keeps its original question bank and written lessons free so you can study even if a paid course is out of reach.

02
Free should mean free.

The original question bank, the worked solutions, the topic lessons, mixed practice, and your readiness score stay free, forever. Full access needs a free account, and practice usage limits apply. Fellow adds AI grading on supported exams, spaced-repetition flashcards, full-length mock exams, analytics, and a study plan that adapts to your progress.

Free forever

Put the formulas to work.

Every formula on this sheet shows up in the free Exam PA question bank, and every question carries a step-by-step solution.

Practice Exam PA questions free →

No credit card. No trial clock.