Using AI to Reduce Assessment Development Time for Educators
Where the hours actually go when building an assessment, and which of them a machine can take.
Ask a teacher how long it takes to build an assessment and the answer is usually vague, because the work is not done in one sitting. It is done in fragments — a free period, an evening, the half hour before a department meeting — and the fragments are hard to add up. That vagueness is part of why the cost is tolerated. Nobody has ever seen the total.
It is worth breaking down, because the parts behave very differently under automation.
The actual breakdown
Deciding what to assess. Short, and it is the professional core of the job. A teacher looks at what was taught, what the objectives were, and what the class seemed shaky on, and decides what the instrument needs to cover. This is fast because it draws on knowledge that already exists.
Drafting the items. Long. Not because any single question is difficult, but because a usable set needs coverage across objectives, a spread of difficulty, variation in form, and no accidental duplication or leakage between questions. Writing twenty items that collectively do a job is a different task from writing twenty items.
Writing the mark scheme. Longer than expected, and usually done last, which is the source of a great deal of trouble. A rubric written after the questions is a reconstruction — it describes what the question seems to want, rather than the questions having been built against a standard. This is where inconsistency between markers originates.
Formatting and administration. Tedious and irreducible by skill. Laying it out, getting it into the right system, producing a version for the students with access arrangements.
Marking. The largest single block in most subjects.
Transcription. Moving marks between the script, the spreadsheet and the system of record. Pure overhead.
Analysis. The reason for doing any of it. Almost always the piece that gets cut, because it comes last and the term has moved on.
Which parts move
Drafting moves substantially. Point the engine at a chapter, a syllabus or a competency framework and it produces items across question types and difficulty bands, each tied to an objective. The teacher's time shifts to reviewing that draft, and reviewing is much faster than composing — recognising that an item tests reading rather than the subject takes seconds, while writing the replacement would have taken ten minutes.
The mark scheme moves, and this is the underrated one. The rubric is drafted alongside the questions rather than afterwards, because a question without a scoring standard is not an assessment. Beyond the time saved, it fixes the sequencing error: the standard and the instrument are now built together, which is the condition under which marking stays consistent across a department.
Formatting and transcription largely disappear. The platform is API-first with native output for the systems your institution already runs, so results move rather than being retyped.
Marking moves for objective and short structured responses, which is most of the volume in most subjects. Extended writing is scored against the rubric as a first pass, with the teacher reviewing rather than originating each judgement.
Which parts do not
Deciding what to assess does not move, and it should not. The engine generates against the objectives it is given; choosing them is the job.
Approval does not move. Every item goes to a teacher to accept or reject, and nothing reaches a learner unapproved. In the first term this is real work — a meaningful share of the time saved on drafting goes back into review. It gets lighter as the approved bank accumulates term over term, and items that fail to separate strong learners from weak ones are retired automatically, so the bank sharpens rather than merely growing.
Anyone quoting a time saving that ignores the review step is quoting a number from a system where nobody has agreed to what the students are being asked.
The compounding effect
The first assessment built this way is faster, but not dramatically so, because review is unfamiliar and the bank is empty. The change is cumulative rather than immediate.
By the second year, a department teaching the same course is drawing on a body of items its own faculty approved, against rubrics it wrote, mapped to its own objectives. New assessments are increasingly assembled from that bank rather than generated from scratch, and generation fills the gaps. The teacher who leaves does not take the bank with them, which is a quiet institutional benefit that rarely appears in a business case.
What to do with the time
This is the part worth being deliberate about, because reclaimed time does not automatically go anywhere useful — it gets absorbed.
The highest-value destination is the analysis step that was previously being cut. When results arrive by learner, cohort and objective while the term is still running, the loop that formative assessment always promised finally closes: eleven students missed the same idea, and Wednesday's lesson can address it.
The second is the small number of scripts where a teacher's judgement genuinely changes an outcome — the ambitious answer, the student whose confusion is interesting, the piece of work that deserves a conversation rather than a mark.
Neither of those is a saving in the accounting sense. Both are the reason the saving is worth having.
AptitudAI — assessment that measures the learner, not the room.

