ExaminaAccuracy Report

AI Grading Accuracy

Accuracy here means agreement with a qualified IB Mathematics examiner, measured at the individual mark level (M1, A1, R1…) rather than on the total score — a grader that reaches 5/7 by awarding the wrong marks has not agreed with anybody. The evaluation set is hand-marked by that examiner and is held back from grading prompt development.

No results yet

We have not published accuracy figures, because we do not have any we would stand behind. The evaluation set is being written by a qualified IB Mathematics examiner. When it is marked and the pipeline has been run against it, the numbers appear here — whatever they say.

What will be measured, and the bar

Mark agreement rate AI and examiner agree on the same individual mark (M1, A1…) ≥ 85%
False positive rate Incorrect work awarded an accuracy mark < 5%
False negative rate Correct work denied an accuracy mark < 10%
Method mark accuracy Agreement on M-marks specifically ≥ 88%

Methodology

Each case is one question with a hand-authored mark scheme and a student response marked decision by decision by a qualified IB Mathematics examiner. Responses span the range that matters — correct, partially correct, and wrong — because a grader that is too generous and one that is too harsh fail in opposite directions, and a set of correct answers would catch neither.

Agreement is measured per individual mark (M1, A1, R1…), not per total score. A result where the AI gives 4/6 and the examiner gives 5/6 for the same reason is a partial disagreement, not a binary fail.

Every published run states the number of cases behind it. A small set can show that something is badly wrong; it cannot show that everything is right, and a percentage without its sample size is marketing rather than evidence.