Every one of FreeFellow's 45,000+ practice questions starts as an AI draft. Before it reaches you it has to pass automated checks, a set of statistical checks that run like software tests, and review by credentialed professionals. Once it's live, it gets recalibrated against real candidate answers when there are enough of them, with at least thirty responses per question. This post walks through each step and the problems each one was built to catch.
I've written separately about the seven tells of an unaudited AI question bank, and most of those came from FreeFellow's own early banks. This post covers what the process looks like now.
Before a question goes live
An AI model drafts each question for a specific learning objective, topic and difficulty target, following a long list of format rules that came from past mistakes. Choices give the answer without arguing for it. Wrong numeric answers come from real mistakes, not numbers spaced evenly around the right one. Choice lengths stay balanced. Each exam's mix of calculation and concept questions is respected. Drafting this way costs a small fraction of an expert-written question, which is why the whole bank can be free.
Next, automated validators check the mechanical things: structure and formatting, math rendering, whether the answer key matches the solution, whether the solution is complete, and dozens of format rules. A question that fails any of them never gets to a reviewer.
Then come the statistical checks. A question can pass every mechanical check and still be gameable, so a set of detectors runs on every content change, and a failure blocks the release the same way a failing test blocks a software deploy. They check:
- Answer-position balance, across the whole bank, by difficulty level and by topic, with hard limits. The by-topic check exists because a candidate noticed a letter pattern that the bank-wide numbers didn't show.
- Giveaway choices: choices that share a leading number and add "because" reasoning, choices that argue for themselves, and sets where one choice is so much longer than the others that its length gives it away.
- Near-duplicates across the whole bank, so the bank's size is real and your score isn't inflated by answering the same question twice.
- Each exam's mix of calculation and concept questions, kept within a target range based on the exam body's outline.
Last, people review it. Questions are written and reviewed under professionals holding the FSA, CFA, CPA, CFP and CAIA. Reviewers look at what the statistics can't: whether the question tests the learning objective or some side detail, whether a wrong choice is clearly wrong or could be argued as right, and whether the solution teaches or just states the answer. Reviewer time is the scarcest part of the process, and the earlier steps exist so it goes where it's needed.
After it goes live
When a question is published, FreeFellow saves a fingerprint of its content. Any later change, like a corrected solution, a rebalanced set of choices or a clearer question, saves the old version instead of overwriting it. Questions with problems that can't be fixed get retired, not deleted, so your past practice records stay the same.
Difficulty keeps getting checked. FreeFellow had recorded more than 870,000 candidate answers as of September 7, 2026. A reviewed calibration pass, with at least thirty responses per question, compares each question's difficulty label with how often candidates got it right on the first try, leaves out outlier accounts that would skew the data, and changes the label where the evidence is strong. One pass moved nearly 1,000 labels. Those measured ratings are what the readiness score is built on.
Candidates can report a suspected error on any question, and a person reviews every report against the source material. Confirmed errors get fixed in the live bank with the old version saved, and the person who reported it hears back. If the question holds up, they get the reasoning. Some of the most useful checks in the process started as candidate reports, and there are probably more blind spots someone will find.
What this means for you
Guessing a letter won't help, since answer positions are balanced at every level. When you pick a wrong choice, its note tells you which mistake it came from. Near-duplicates and giveaway choices are filtered out, so your practice score reflects what you know. And the difficulty behind your readiness score comes from how real candidates did, not from someone's guess.
Anyone can try the sample questions without signing up, and the full bank needs a free account. You can run the checks from this post on FreeFellow yourself.