ExaminaAccuracy Report

AI Grading Accuracy

Examina is evaluated against a sealed set of 200 IB-style questions authored and hand-marked by a qualified IB Mathematics examiner. The AI is never trained on this set. Agreement is measured at the individual mark level (M1, A1, R1…) — not just the total score.

🔬

Evaluation in progress

We are running the grading pipeline against a sealed 200-question gold set. Results will be published here once the mark agreement rate exceeds 85% on two consecutive nightly runs.

What we measure

Mark agreement rate AI and gold examiner agree per individual mark (M1, A1…) ≥ 85%
False positive rate Incorrect work awarded an accuracy mark < 5%
False negative rate Correct work denied an accuracy mark < 10%
Method mark accuracy Agreement on M-marks specifically ≥ 88%

Methodology

The gold set consists of 200 original IB-style questions authored by a qualified IB Mathematics examiner. Each question has a hand-authored mark scheme and three representative student responses (poor, partial, correct) — 600 marked submissions in total. This set is sealed and never used in grading prompt development.

Agreement is measured per individual mark (M1, A1, R1…), not per total score. A result where the AI gives 4/6 and the gold gives 5/6 for the same reason is a partial disagreement, not a binary fail. Results are published only when the mark agreement rate exceeds 85% on two consecutive nightly runs.