Can AI Grade Your Essay Exam Answers? An Honest Look
AI grading works under exactly one condition: the rubric. Grounded in a real grading standard, a frontier model is the closest thing to a personal grader most candidates will ever have, which is why FreeFellow reproduces every released SOA Exam PA sitting and all 13 released CAS Exam 5 sittings as graded walkthroughs and grades candidate answers against rubrics derived from published solutions and grader commentary. Ungrounded, the same model becomes a flattery machine that scores fluent answers well and responsive answers no better.
I am Jeffrey Ting, FSA, CFA, the founder of FreeFellow, and constructed response is where I think AI genuinely changes exam prep, so this post is both enthusiastic and specific about the failure mode.
The Feedback Desert
Multiple choice grades itself. Essay and written-answer formats, CFA Level III essays, SOA Exam PA, CAS written-answer papers, CPA simulations, historically had two options: pay a human grading service real money for feedback on a handful of essays, or practice with no feedback at all. Most candidates chose the second, which means practicing blind on the exact format that fails the most candidates. You cannot see your own unresponsiveness; that is precisely the defect a grader sees instantly.
What AI Grading Does Well
Given the question, the rubric, and your answer, a frontier model applies the standard with a consistency human graders work hard to match. It reads your answer at midnight with the same attention as at noon. It grades your twentieth practice essay with the same patience as your first. And it is specific in the way that makes feedback usable: which rubric points you earned, which you missed, and what sentence would have earned the missed ones.
Speed matters more than it sounds. Feedback that arrives in seconds, while your reasoning is still in working memory, changes behavior; feedback that arrives in two weeks from a grading service is an autopsy.
The Failure Mode: Flattery
Ask a bare chatbot to "grade my essay answer" and it anchors on fluency and plausibility. Language models are trained to be agreeable, and an ungrounded grader drifts generous, scoring well-written answers that never respond to the command word. Exam graders do the opposite: they hunt for specific, rubric-listed items and award nothing for eloquence.
The fix is structural, not clever prompting: supply the actual grading standard and instruct the model to award points only for rubric items explicitly present in the answer, citing the sentence that earns each point. Grounded that way, the generosity mostly disappears, and what remains is the honest gap between what you wrote and what the standard rewards.
Where the Rubric Comes From
This is the part that cannot be improvised in a chat window. A usable rubric has to reflect how the exam actually awards points, which means deriving it from the released material: published solutions, sample graded answers, and grader commentary. For the actuarial written-answer exams, the released sittings make this concrete: FreeFellow reproduces every released Exam PA sitting and all 13 released CAS Exam 5 sittings as graded walkthroughs, each with the rubric, the published sample answers and model solutions, and the grader commentary that explains what full credit required. CFA Level III essay practice follows the same pattern with original questions built to the published grading style.
How to Do This Free
FreeFellow's free tier ships a copy-to-AI prompt builder on constructed-response practice: it assembles the question, the rubric, and the model solution into a single prompt, so you can paste your answer into your own ChatGPT or Claude and get rubric-grounded grading at $0. The paid Fellow tier does the same thing instantly in-app, graded by Claude against the same rubric, with uncapped grading on the annual tier.
The workflow that compounds: write your answer under exam-format time pressure first, before seeing any solution. Get the graded feedback. Rewrite the answer to earn the missed points. Then read the grader commentary and note what the exam rewards that you did not predict. Candidates who skip the rewrite step lose most of the value; the rewrite is where the rubric's logic becomes yours.
What AI Grading Cannot Tell You
Honesty section. It practices the writing, not the sitting: no AI feedback replicates exam-day time pressure, handwriting fatigue on paper-based exams, or the exam body's actual grading decisions in a given administration. Scores from any practice grading, AI or human, are direction, not prediction. Nobody grades the real thing but the exam body, and no preparation tool should imply otherwise.
Used inside those limits, rubric-grounded AI grading removes the oldest inequity in essay-format prep: feedback used to be a luxury good. It is not anymore. Start with the free constructed-response practice, write badly, get graded, rewrite, and let the rubric teach you what the exam actually buys.