← Back to blog

2026-06-14 • Partnerships • Natasha

426 reports at university scale: our IS1108 pilot with NUS Computing

!IS1108 Ethics in Computing — cohort report

Feedback at university scale usually forces a trade-off between depth and speed. You can give each student detailed feedback, or you can return the work quickly across a large cohort. You rarely achieve both.

Last semester Ren ran its largest university pilot to date, with NUS Computing's IS1108. Ren marked individual reports for all 426 students.

Feedback grounded in how the module is taught

We built a dedicated IS1108 context layer from the module's own materials: its FISh framework, its marking criteria, and its teaching approach. The goal was simple. Feedback must not be generic. It must reflect how the teaching team taught IS1108.

Every student received detailed written feedback tied to the module's scope and expectations. The teaching assistants produced, reviewed, and approved work that would have taken far longer by hand, and the standard stayed more consistent.

!Teaching assistants mark IS1108 reports in Ren

The teaching assistants worked inside Ren. They read each report next to the AI's feedback, tags, and topic map. They then edited the feedback and approved it before any student saw it.

Why consistency matters

Marking is not only a question of speed. It is a question of fairness.

Ren produces feedback from one shared context layer and reviews it against one set of marking criteria. Every student therefore meets a more consistent standard. That standard is harder to hold when the work spreads across many markers, because each marker carries their own reading of the criteria.

We measured this directly. With 19 markers on the panel, scores on the same report spread across 4.5 marks on a 20-mark scale. Ren's marks deviated from the approved scores by 0.40 marks on average — a tighter band than even a two-marker panel typically achieves.

The numbers

We benchmarked Ren's suggested marks against the scores the teaching assistants reviewed and approved, across 390 reports on a 20-mark scale.

| Metric | Result |

|--------|--------|

| Mean absolute error | 0.40 marks |

| Within 1 mark | 95% |

| Within 2 marks | 100% |

| Human spread on the same paper (19 markers) | 4.5 marks |

Speed moved just as much as consistency. Marking a report by hand took the teaching assistants around 33 minutes. Reviewing and approving Ren's feedback took around 3.3 minutes — a 10× reduction, which across the cohort returned roughly 250 hours of marking time to the teaching team.

Insights that usually stay buried

After the batch, we gave cohort-level analytics to Prof Boon Kee Lee. The analytics show where students struggled, which questions caused the most confusion, the most common misconceptions, and an estimate of the time saved.

Manual marking usually hides these insights. Nobody sees them, even after hundreds of hours of work.

Why we built Ren

This pilot reminded us why we started. We want to help educators give better feedback at scale, and never at the cost of quality.

Our thanks go to Prof Boon Kee Lee, one of our earliest supporters, and to all 19 teaching assistants who made the pilot possible.


We have a few pilot places open for next semester. If you want faster, fairer, and more consistent feedback, talk to our team.