Why we benchmark
Claims about AI accuracy are easy to make and hard to verify. For that reason we benchmark Ren against experienced human markers, and we publish the results.
Methodology
Dataset
- A mix of short-answer questions (1 to 4 marks) and extended responses (6 marks or more).
- Eight topics across the chemistry O-Level syllabus.
Marking process
Two markers scored each response:
- Human markers — experienced GCSE teachers.
- Ren's AI marking system — with the answer scheme uploaded.
We treated the teacher scores as ground truth and measured Ren's agreement against them. Note the limit of this design. It measures how closely the AI matches human marking. A mismatch can have several causes: human error, AI error, or a vague answer scheme.
Results
| Metric | Ren score |
|--------|-----------|
| Exact mark agreement | 82.4% |
| Within 1 mark | 96.1% |
| Cohen's Kappa (against consensus) | 0.79 |
| Mean absolute error | 0.31 marks |
For context, two human markers agree exactly at about this rate. Ren therefore performs inside the range of normal human marker variation.
By question type
| Question type | Exact agreement | Within 1 mark |
|---------------|-----------------|---------------|
| Short answer (1 to 2 marks) | 91.2% | 99.1% |
| Medium response (3 to 4 marks) | 79.8% | 95.3% |
| Extended response (6 marks or more) | 68.4% | 89.2% |
Accuracy is highest on structured short-answer questions, as we expected. It falls on longer and more subjective responses. Human markers behave the same way. Agreement between two teachers also falls on extended responses.
Key takeaways
What Ren does well
- Factual accuracy. Ren identifies reliably whether a student included the required scientific facts and key terms.
- Structure recognition. The model identifies reliably whether a response follows the expected structure, such as a "describe and explain" format.
- Consistency. Ren does not become tired. It marks the 500th paper with the same attention as the first.
Where Ren needs teacher review
- Borderline cases. A response that sits on a grade boundary needs human judgement.
- Unusual answers. A student can show understanding through an unconventional approach. A teacher must confirm that the reasoning is valid, especially when the answer scheme does not describe it.
- Handwriting artefacts. On a scanned response, very poor handwriting or an unclear drawing reduces accuracy.
Continuous improvement
We repeat these benchmarks each quarter and publish the new results. Our research and engineering team improves the marking engine three ways:
- We expand the data set with material from partner schools.
- We improve how the AI grounds its marks in the answer scheme.
- We work with teachers on the edge cases that earlier benchmarks identified.
Coming on board
We run a benchmark for each early partner. This confirms that the product works for your papers before you depend on it.
Do you want to run your own benchmark? Get in touch and we will set it up.