Why AI Will Not Replace Teachers but Will Transform Assessment Workflows
The work worth automating is the marking, the drafting and the admin. The judgement stays where it was.
The replacement question is usually asked at the wrong level. It treats teaching as one job that either survives automation or does not, when it is a bundle of quite different activities that happen to be performed by the same person. Some of them are professional judgement. Some of them are clerical work that accumulated around the judgement because there was nobody else to do it.
Separating the two is more useful than arguing about the bundle.
The parts, listed honestly
Automatable now, without much argument:
- Drafting first-pass items from material that already exists
- Applying a written scoring standard to objective and short-answer responses
- Tallying, transcribing and moving results into the system of record
- Identifying which questions a cohort found hard, and which students share a gap
Not automatable, and not delegable:
- Deciding what is worth assessing in the first place
- Judging whether a question is fair to this class, with this history, this week
- Recognising that a wrong answer is a step forward from the last wrong answer
- Reading extended work and knowing that a fifteen-year-old has attempted something genuinely ambitious
- Deciding what to do on Monday
The first list is where the hours go. The second list is the profession. Almost every anxiety about this technology is really an anxiety that a vendor will conflate them — and that anxiety is well earned, because some do.
Why the drafting can move but the approval cannot
Generation shifts the teacher from author to editor. That is a real reduction in effort, and the reason is not that judgement has been removed: it is that recognising a bad question is much faster than composing a good one.
A teacher who knows their class can see in seconds that an item tests reading rather than the subject, or leaks its answer in the phrasing, or sits in a context half the room will not recognise. Rejecting it costs a click. Writing the replacement would have cost ten minutes. The review stage exists because that judgement is the part that cannot move — nothing reaches a learner unapproved, and the surviving set becomes the institution's bank.
The failure mode is obvious and worth naming. If approval degrades into bulk acceptance under time pressure, nothing has been saved. The cost has been deferred to the lesson where thirty students hit a broken question at once, and then to the parent email. A workflow that makes approval fast is doing its job; one that makes approval feel optional has removed the only safeguard in the system.
Where automated marking stops
Objective items and short structured answers mark reliably against a rubric. That is most of the volume in most subjects, and reclaiming it is the difference between analysis happening on the evening of the test and not happening at all.
Extended writing is different, and the honest position is that it is different in kind rather than in difficulty. A model can check whether an essay addresses the stated criteria, and that is genuinely useful as a first pass. It cannot tell you whether an argument is interesting, whether a risk paid off, or whether this particular student has just done the best work of their year. Those judgements are comparative, contextual and personal — they depend on knowing the writer.
The workable arrangement is therefore not "AI marks everything" or "AI marks nothing". Routine responses are scored automatically, extended work is drafted and flagged, and the teacher spends the reclaimed hours on the scripts where their judgement changes the outcome. That is a better use of a professional than tallying.
The workflow that actually changes
Formative assessment has never failed on theory. Assess, find the gap, act before the unit ends is uncontroversial and has been for decades. It fails on turnaround, in the same place every time: marking finishes late, analysis does not happen, the next unit starts, and the gap compounds until it surfaces in the summer.
What changes when results land by learner, cohort and objective on the day is that the last step survives contact with a real week. Eleven students who missed the same idea are a visible pattern on Tuesday evening rather than a vague impression a fortnight later. Ten minutes of reteaching goes into Wednesday's lesson.
That is the transformation. Not a different profession — the same profession, with the loop closing often enough to be worth having.
What a teacher should demand of any such system
- A real gate. You can reject any item, the rejection is recorded, and nothing reaches a student without a name against it.
- A written rubric per item, before the class sits it. If the scoring standard is not written down, marking drifts between the first script and the thirtieth, and feedback stays a number.
- Results by objective, not just by total. A percentage is a fact about a student. A missing sub-skill is an instruction.
- Practice kept separate from assessment. If students can practise their way into seeing the exam, the exam measures nothing. The separation has to be enforced by the system that generates both.
- Your bank stays yours. Those approved items were authored by your department's judgement. Establish that you can take them.
The part worth watching
The genuine risk in this technology is not redundancy. It is deskilling by convenience — a teacher who stops writing rubrics because the machine drafts them, stops noticing why an item is weak because most arrive serviceable, and gradually loses the practised judgement that made them a good gate in the first place.
That is not inevitable, but it is not automatic to avoid either. It is the reason the approval step should stay slightly effortful, and the reason a department should keep talking about why an item was rejected rather than only recording that it was.
AptitudAI — assessment that measures the learner, not the room.

