How AI Can Transform India's Entrance Exam Culture
A single sitting decides millions of futures. Adaptive measurement changes what that sitting has to carry.
The Indian entrance exam is one of the most consequential instruments any education system has ever built. A few hours on a particular morning determine which institution a student attends, which profession is available to them, and in a great many families, what the next twenty years look like. The scale is extraordinary and so is the seriousness with which it is taken.
The criticism usually levelled at it — that it is stressful — is true but shallow. Stress is a symptom. The structural problem is that a single sitting is being asked to carry more information than a single sitting can hold, and everything distorted about the culture around it follows from that overload.
What the format forces
When one number allocates scarce places among millions of candidates, three things become inevitable.
The exam must discriminate finely at the top. When lakhs of students compete for a few thousand seats, the difference between rank 900 and rank 4,000 has to come from somewhere, and it comes from a handful of unusually hard items. The result is an instrument optimised for separating the very strongest candidates, which measures everyone else imprecisely almost as a by-product.
The exam must be identical for everyone, for reasons of fairness that are entirely legitimate. But identical papers mean fixed difficulty, and fixed difficulty means most candidates spend most of their time on questions that tell the system little about them.
And because the paper is fixed and predictable in form, preparing for the paper becomes a separate discipline from learning the subject. This is the origin of the coaching economy. Coaching is a rational response to a measurable fact: familiarity with the instrument raises the score independently of subject mastery. As long as that is true, families who can pay will pay, and the exam will partly measure ability to pay.
Where adaptive measurement changes the arithmetic
An adaptive test does not hand every candidate the same paper. Each response determines the next item, so the sequence converges towards the boundary of what a particular student can do. Two consequences follow that matter specifically for this system.
The first is precision across the whole range rather than at one end. A fixed paper is accurate near its own difficulty and vague elsewhere. An adaptive instrument spends its questions where they carry information for that candidate, which means a student in the middle of the distribution is measured about as precisely as one at the top. In a system that allocates on rank, precision across the range is not a technical nicety. It determines whether the ranking is real.
The second is that a predictable paper becomes harder to drill. When items are generated rather than drawn from a fixed and eventually memorised pool, and when the sequence differs by candidate, preparation that consists of pattern-recognition on past papers loses much of its edge. This does not abolish coaching, and nobody should claim it does. It shifts what coaching is for, from familiarity with an instrument towards the subject itself.
Language, which is usually treated as a footnote
A student sitting an entrance exam in a language that is not the one they learned the subject in is being tested on two things and credited for one. In a country where instruction runs in many languages and the high-stakes exams have historically run in few, this is not a marginal effect.
Multilingual generation is one of the more consequential capabilities here and one of the least discussed. Generating and scoring in the languages a programme actually runs in, rather than translating a paper written in English after the fact, changes what the score means. Translation after the fact tends to preserve the words and lose the difficulty calibration; an item that was straightforward in one language can become a reading comprehension problem in another.
The related point is bias in the items themselves. Questions carry contextual load — a word problem set in a milieu one group of candidates knows well and another has never encountered is testing that milieu. Checking items for linguistic, cultural and contextual load before they reach a candidate, and flagging the uncertain ones for human review, does not eliminate this. It makes it visible and auditable, which authoring by hand at national scale never can be.
What continuous measurement would actually change
The deeper opportunity is not a better entrance exam. It is reducing how much the entrance exam has to carry.
At present the single sitting is the only reliable signal about a candidate, because school assessment varies too much across boards, states and institutions to be comparable. If continuous, curriculum-aligned assessment ran through the school years — the same instrument, the same standard, results by objective rather than by aggregate — a selection system would have a longitudinal record to draw on, and a much richer one. A student who improved steadily over three years is a different proposition from one who peaked on a Tuesday, and no single sitting can tell them apart.
That is a policy change, not a technology purchase, and it would take years. But the technical obstacle that historically blocked it, the cost of generating and marking assessment at population scale, is the part that has genuinely moved.
The cautions
Adaptive testing at this scale raises real questions, and the ones worth taking seriously are not about the algorithm.
Candidates receiving different item sequences must be scored on one defensible scale, and the method has to be explainable to a family that wants to contest a result. Opacity in a high-stakes system is corrosive regardless of accuracy.
Delivery has to work where candidates actually are. An adaptive system that assumes reliable connectivity and a personal device would exclude precisely the students the fairness argument is meant to protect.
And a longitudinal record of a child's performance is a serious thing to create. Used to target support early, it is valuable. Used to sort children earlier than the current system does, it would be worse than what it replaced. That is a governance decision, and it should be made deliberately rather than arriving as a side effect of better measurement.
The prize is worth the care. Not the abolition of a demanding exam culture — a demanding system is not the problem — but a version of it where what is being measured is the subject, and where the measurement is precise enough to deserve the weight it carries.
AptitudAI — assessment that measures the learner, not the room.

