Yes, ChatGPT can write a multiple-choice question that looks right in a few seconds. I still wouldn't study from AI questions that nobody has checked. I've seen what goes wrong up close, because FreeFellow's first drafts are written by AI too, and its checks have caught every problem in this post in its own questions. One early bank had 47% of its correct answers on the same letter. A later sweep found more than 2,000 answer choices that explained why they were right, which real exams never do.
I'm Jeffrey Ting, FSA, CFA, and I run FreeFellow. I don't have a problem with AI writing the first draft. The problem is when nobody reviews it. Here are seven signs that nobody did, and you can check a bank for every one of them yourself.
The seven tells
1. The right answer keeps landing on the same letter
AI usually writes the correct answer first and the wrong ones after, so unless a separate step shuffles them, the right answer ends up on A a lot. One of FreeFellow's banks was evenly spread for months. Then two new batches went in without the shuffling step, and the bank drifted to 47% A. Someone who always guessed A would have scored 47% without reading a single question.
It can also be hard to spot. A candidate once wrote in that most of the right answers in one bank were A. The bank as a whole was 25% on each letter, so it looked fine, but she was partly right: one topic inside it was 33% A. FreeFellow now checks the answer spread on every content change, by topic and by difficulty level, with hard limits (no letter above about 1.7 times its fair share, none below 0.4 times). You can do a rough version yourself. Count the correct letters over 40 questions. On a four-choice exam, anything near 40% on one letter isn't chance.
2. On math questions, the middle number is usually right
Ask a model for a calculation question with five numeric choices and it'll work out the right answer, then put the wrong answers around it, a bit lower and a bit higher. Actuarial and finance exams list numeric choices from smallest to largest, so the right answer ends up in the middle. One of FreeFellow's early actuarial banks had it there 47% of the time.
On released SOA and CAS exams, the right answer is the smallest or largest choice about a third of the time. That's because real wrong answers come from real mistakes, like forgetting to annualize, dropping a sign or using the wrong table, and those mistakes tend to land on one side of the right number. FreeFellow rebuilt its numeric wrong answers from mistakes like these, which helped the solutions too: each wrong choice now matches a mistake candidates actually make, and the solution can tell you which one. To check a bank, count how often the right answer is the smallest or largest choice across 30 calculation questions. If it's almost never, a machine built the wrong answers.
3. Answer choices that explain themselves
Real exam choices give an answer, not a reason. Unchecked AI questions have choices like "527,000, because the translation loss is excluded from net income" next to "505,000, because unrealized gains reduce comprehensive income." The reasoning belongs in your head during the exam and in the solution afterward.
It's more than a style problem. When a model can't come up with wrong answers that differ in the numbers, it leans on the explanation to tell them apart, and the explanation often covers up bad math. In one FreeFellow audit of 39 of these questions, 16 had at least one "because" whose math didn't come out to the number it stated. One cleanup pass removed the reasoning from about 1,900 choices, and three checks now block any new ones. Where removing it would leave two identical choices, the question gets rewritten instead. To check a bank, search the choices for ", because". You won't find it on a real exam.
4. The longest answer is right
Test writers have known about this one for a long time. The right answer ends up longest because the author puts the most detail into it. AI does this a lot, and if the most detailed choice keeps being right, you can score well without knowing the material. FreeFellow flags questions where one choice is much longer than the others and sends them for a rewrite.
5. The same question, reworded
Ask a model for 1,000 questions on one syllabus and it keeps coming back to the same setups: the same bond, the same facts, a few words changed. Near-duplicates make a bank look bigger than it is, and they push your practice score up because you're answering questions you've already seen. FreeFellow checks the whole bank for similar questions and rewrites the twins to test the learning objective a different way. One campaign replaced 121 near-duplicate pairs across 23 exams, and the check now runs on all new content.
6. Numbers from the wrong year
A model learns from years of text, so it mixes the years up. Ask it for a tax question and it might use last year's thresholds, an old lease standard or an old syllabus's topic weights, and it won't mention it. On tax-heavy exams like the EA, CPA REG and TCP, and the CFP, one outdated number changes the right answer. FreeFellow ties its tax figures to the current law and each exam's testing window, keeps those banks out of bulk find-and-replace edits, and rewrites affected questions with the whole question in view so any math that depends on the number changes with it. It also watches the exam bodies' own pages for syllabus and policy changes. If you're using AI-written questions from somewhere else, ask which tax year the numbers assume, then check one against the IRS or the exam body.
7. The wrong kind of question
Every exam has its own mix of calculation and judgment questions, and AI tends to write whatever's easiest, which is usually calculation. FreeFellow's FRM Part 2 bank launched at about 48% calculation questions, when the real Part 2 is mostly qualitative, at roughly 30 to 35% calculation. A candidate called it out, and he was right. Every FreeFellow exam now has a target range for its calculation mix, based on the exam body's outline and what candidates report, and the same checks enforce it. The FRM Part 2 bank was rebalanced with caps by topic to match the real exam.
Why difficulty has to be measured
Even a perfect question writer can't tell you how hard a question is. Difficulty comes down to how many prepared candidates get it right, so when a chatbot labels its own question "hard," it's guessing. FreeFellow measures it instead, from more than 870,000 candidate answers as of September 7, 2026. Difficulty gets recalibrated in a reviewed pass that needs at least thirty responses per question. It compares each question's label with how often candidates got it right on the first try and changes the label where the evidence is strong. One pass moved nearly 1,000 labels. Those measured ratings feed the readiness score.
A ten-minute check for any question bank
Before you trust a question bank, AI-written or not, check these. Count the correct letter across 40 questions and look for anything near 40% on one letter. Count how often the right numeric answer is the smallest or largest choice, and be suspicious if it's close to never. Search the choices for ", because". See whether the longest choice keeps winning. Then ask when the difficulty ratings were last updated and what data went into them. If nobody can tell you, that's a bad sign. FreeFellow has open sample questions, and the full bank is free with an account, so you can run these checks on it too.
None of this means you shouldn't study with AI. A chatbot is great for explanations. Use it to rephrase a concept until it clicks, to walk through a solution you don't follow, or to quiz you on a topic you pick. FreeFellow's free tier even has a copy-to-AI prompt builder for written-answer practice. Just make sure the questions you measure yourself against have been checked by someone.