AI Grading Accuracy
Examina is evaluated against a sealed set of 200 IB-style questions authored and hand-marked by a qualified IB Mathematics examiner. The AI is never trained on this set. Agreement is measured at the individual mark level (M1, A1, R1…) — not just the total score.
Evaluation in progress
We are running the grading pipeline against a sealed 200-question gold set. Results will be published here once the mark agreement rate exceeds 85% on two consecutive nightly runs.
What we measure
Methodology
The gold set consists of 200 original IB-style questions authored by a qualified IB Mathematics examiner. Each question has a hand-authored mark scheme and three representative student responses (poor, partial, correct) — 600 marked submissions in total. This set is sealed and never used in grading prompt development.
Agreement is measured per individual mark (M1, A1, R1…), not per total score. A result where the AI gives 4/6 and the gold gives 5/6 for the same reason is a partial disagreement, not a binary fail. Results are published only when the mark agreement rate exceeds 85% on two consecutive nightly runs.