Skip to content
AptitudAI
Policy

Beyond High-Stakes Testing: How AI Is Redefining Assessment for the Future of Learning

One exam at the end of the year decides a lot and explains very little. What replaces it when measurement can run continuously.

High-stakes testing is usually defended on fairness and attacked on stress, and both arguments miss the more basic problem. A terminal exam is a single observation used to support several claims at once: what a student knows, how reliably they know it, how they compare with a cohort, and how they are likely to do next. One observation cannot support four claims. It supports one, with a wide error bar, and the system treats the error bar as if it were not there.

Everything downstream follows from that overload. Teaching narrows towards the instrument because the instrument carries all the weight. Preparation becomes a discipline separate from the subject. And the result arrives after the point at which anyone could have used it, which is why a terminal exam is an excellent accountability tool and a nearly useless teaching one.

What high-stakes testing is genuinely good at

Reform arguments tend to skip this, and then fail on contact with the people responsible for standards.

A common paper, sat under supervision, marked to one scheme, is cheap at scale, hard to corrupt, and legible to outsiders. An employer or an admissions office can trust it without knowing anything about the school that prepared the candidate. In a system where institutional quality varies and trust between institutions is thin, that portability is the entire function. Continuous internal assessment, whatever its pedagogical advantages, does not have it: it asks a third party to trust a judgement made by someone with an interest in the outcome.

Any serious proposal has to preserve portability. Most do not, which is why most go nowhere.

The distinction that actually matters

The useful split is not between testing and not testing. It is between measurement used to decide and measurement used to teach, and the two have almost opposite requirements.

Deciding needs comparability, defensibility, security and a fixed point in time. It should be rare, formal and invigilated, and it is reasonable for it to be uncomfortable.

Teaching needs frequency, granularity, immediacy and low consequences. It should be routine, objective-level, returned the same day, and safe to fail — because a measurement that a learner is afraid of produces defensive behaviour rather than information.

Trying to make one instrument do both is the original error. The terminal exam has been asked to serve teaching, at which it is hopeless, and continuous coursework has been asked to serve certification, at which it is contested. Separating them is not a compromise between the two positions; it is the design that both positions were reaching for.

What changed to make the split practical

The reason nobody separated them before is that the teaching half was unaffordable. Frequent, objective-mapped, consistently marked assessment across every class in an institution is an enormous amount of expert authoring and marking time, and expert time is the scarcest resource in education. So schools ran a few internal tests, marked to varying standards, and relied on the terminal exam for everything else.

Three costs have moved. Items can be drafted from an institution's own material rather than authored from nothing, with faculty approving or rejecting each one. Rubrics are written alongside the questions rather than reconstructed afterwards, which is what makes marking consistent between one teacher and the next. And objective and short-response marking is automatic, which is most of the volume in most subjects.

What that buys is not a replacement for the terminal exam. It is the thing that was always missing beside it: a continuous, comparable record of what students actually know, in time to act on it.

What a policymaker should ask for

  • Objective-level reporting, not aggregates. A grade distribution tells you how a cohort did. Objective-level data tells you which parts of the curriculum are failing everywhere, which is a different problem with a different remedy.
  • Human approval, recorded. Generated items reaching learners without institutional sign-off produce data that is precise and untrustworthy. The audit trail matters more than the generation.
  • Bias checked before delivery, not after complaint. Every item carries linguistic, cultural and contextual load. Checking for it at authoring time is part of the work; discovering it from an appeal is too late for the candidates already affected.
  • Delivery in the languages the system actually runs in. An exam sat in a second language measures two things and reports one number.
  • Separation between practice and assessment, enforced structurally. If students can practise into the exam pool, the exam measures preparation access rather than learning.

The failure modes worth planning against

Continuous assessment becomes continuous high-stakes assessment. This is the most likely way the reform goes wrong. If every fortnightly check is recorded, reported and attached to a student's file, the stakes have not been lowered — they have been distributed, which is worse. Low-stakes measurement has to be genuinely low-stakes, which means some of it must not be kept.

Surveillance replaces trust. A system capable of continuous measurement is also capable of continuous monitoring, of students and of teachers. The moment assessment data enters a performance appraisal, teachers begin setting assessments that produce good data, and the honest picture the system was built to provide is gone within a year.

Sorting happens earlier. A longitudinal record makes it possible to predict outcomes sooner. Used to target support, that is valuable. Used to route children onto pathways at twelve rather than sixteen, it is a considerably worse system than the one it replaced, delivered with better evidence.

None of these are technical risks. They are governance decisions that will be made by default if they are not made deliberately.

What is realistically ahead

Not the abolition of the terminal exam. It does a job nothing else does, and the systems that have tried to remove it have generally reinstated it.

What is available is narrower and more useful: reducing how much that exam has to carry. If a student arrives at it with three years of comparable, objective-level evidence behind them, the exam no longer has to serve as the only reliable signal about them. It can be one observation among many rather than the observation, and a selection process can look at trajectory as well as position — which is a better description of an eighteen-year-old than any single morning can be.

That is a modest ambition compared with most rhetoric about the future of assessment. It is also the version that could actually be implemented without asking anyone to give up the thing they were right to defend.

AptitudAI — assessment that measures the learner, not the room.