AI Grading Accuracy
Accuracy here means agreement with a qualified IB Mathematics examiner, measured at the individual mark level (M1, A1, R1…) rather than on the total score — a grader that reaches 5/7 by awarding the wrong marks has not agreed with anybody. The evaluation set is hand-marked by that examiner and is held back from grading prompt development.
No results yet
We have not published accuracy figures, because we do not have any we would stand behind. The evaluation set is being written by a qualified IB Mathematics examiner. When it is marked and the pipeline has been run against it, the numbers appear here — whatever they say.
What will be measured, and the bar
Methodology
Each case is one question with a hand-authored mark scheme and a student response marked decision by decision by a qualified IB Mathematics examiner. Responses span the range that matters — correct, partially correct, and wrong — because a grader that is too generous and one that is too harsh fail in opposite directions, and a set of correct answers would catch neither.
Agreement is measured per individual mark (M1, A1, R1…), not per total score. A result where the AI gives 4/6 and the examiner gives 5/6 for the same reason is a partial disagreement, not a binary fail.
Every published run states the number of cases behind it. A small set can show that something is badly wrong; it cannot show that everything is right, and a percentage without its sample size is marketing rather than evidence.