How FreeFellow Writes and Audits 40,000+ Questions
Every one of FreeFellow's 40,000+ practice questions starts as an AI draft, and every one must then survive deterministic validators, statistical release gates that run like software tests, review under credentialed professionals, and a monthly recalibration against real candidate answers. This post walks the full pipeline, including the defects it was built to catch, because in a market flooding with unaudited AI content, the audit process is the product.
I am Jeffrey Ting, FSA, CFA, the founder. I have written elsewhere about the seven tells of an unaudited AI question bank, most of them found in FreeFellow's own early banks. This is the companion piece: what the pipeline looks like once those lessons are encoded.
Stage 1: Drafting
A frontier model drafts each question against a specific learning objective, topic, difficulty target, and a long list of format rules learned the hard way: choices state the answer without arguing for it, numeric distractors are derived from real mistake paths rather than symmetric offsets, choice lengths stay balanced, and each exam's calculation-versus-conceptual style is respected. Drafting this way costs a small fraction of an expert-written question, which is why the entire bank can be free.
Stage 2: Deterministic Validation
Before a draft goes anywhere, automated validators check the mechanical facts: structure and formatting, math rendering, answer-key consistency, solution completeness, and dozens of format rules. These checks are boring by design. A question that fails any of them never reaches a human reviewer, let alone a candidate.
Stage 3: Statistical Release Gates
The interesting layer. A bank can be mechanically valid and still be exploitable, so a battery of statistical detectors runs on every content change, and a regression blocks release the way a failing test blocks a software deploy:
- Answer-position balance, checked bank-wide, per difficulty tier, and per topic, with hard thresholds. The per-topic check exists because a candidate correctly sensed a letter pattern that bank-wide statistics could not see.
- Giveaway-choice detection: choices that share a leading value with appended "because" reasoning, choices that argue their own case, and lopsided length spreads that make the longest answer a tell.
- Near-duplicate detection across the corpus, so the bank's size means what it says and your score is not inflated by re-answering the same question in different clothes.
- Calculation-versus-conceptual mix held inside a target band per exam, sourced from each exam body's outline, so the bank tests the way the real exam tests.
Stage 4: Credentialed Human Review
Questions are written and reviewed under professionals holding the FSA, CFA, CPA, CFP, and CAIA designations. Humans own what statistics cannot see: whether the question tests the learning objective or trivia adjacent to it, whether a distractor is defensibly wrong or ambiguously right, and whether the solution teaches or merely asserts. Human attention is the scarcest input in the pipeline, which is exactly why the earlier stages exist: to spend it only where it is needed.
Stage 5: Versioned Publication
When a question ships, its content is fingerprinted. Any later edit, a corrected solution, a rebalanced choice set, a clarified stem, snapshots the prior version rather than overwriting history. Questions with confirmed unfixable defects are retired, not deleted, so past practice records stay intact. Nothing is silently swapped under a candidate's feet.
Stage 6: Live Calibration
Publication is the midpoint. Millions of real candidate answers feed a monthly calibration pass that compares each question's difficulty label against actual first-attempt accuracy, excludes outlier accounts that would poison the data, and re-labels where the evidence is statistically strong. One round moved nearly 1,000 labels. This is the loop no static book and no chat session has, and it is what makes the readiness score an estimate rather than a vibe.
Stage 7: The Report Loop
Candidates report suspected errors from any question, and a person reviews every report against source material. Confirmed errors are fixed in the public bank with the prior version preserved, and the reporter hears back with the resolution. Defensible questions get their reasoning explained. Some of the pipeline's best gates started as candidate reports; the process assumes the next blind spot exists and someone outside will find it first.
What This Buys You
The practical answer, from a candidate's chair: a guessing strategy earns you nothing, because positions are balanced at every slice. Every wrong choice teaches, because distractors encode real mistakes with per-choice notes naming them. Your practice score tracks knowledge, because near-duplicates and giveaway shapes are gated out. And the difficulty behind your readiness number is measured against thousands of real candidates, not asserted.
All of it is browsable free, without signup, which is deliberate: the checks in this post can be run by anyone, on us, before trusting a single question.