Skip to content
AptitudAI
Research

Learning Digital Twins: Modeling a Student's Cognitive Growth with AI

If you can model how a student's understanding moves, you can teach ahead of the gap instead of behind it.

A digital twin, in the sense the term was borrowed from, is a running model of a physical thing — a turbine, an engine, a production line — kept in step with the real object by a stream of measurements, and used to answer questions about the object that would be expensive or destructive to ask directly. What happens if this bearing runs another thousand hours. Where does this line fail first under load.

Applied to a learner, the idea is straightforward to state and much harder to earn: a model of what a student currently understands, updated by everything they do, accurate enough that you can ask it what they will struggle with next.

The reason to be interested is timing. Teaching currently responds to gaps after they surface in a result. A model that predicts where the next gap will open lets the intervention arrive first, which is the difference between remediation and instruction.

What the model is actually made of

It is worth deflating the vocabulary before defending the concept. A learning twin is not a simulation of a mind. It is a structured estimate of mastery over a set of objectives, with uncertainty attached, updated as evidence arrives.

Three things distinguish it from a gradebook, which is also a record of performance.

It is organised by objective rather than by assessment. A gradebook says a student scored 58 on Tuesday. A twin says which of the twelve objectives that paper touched are held, which are not, and how confident the estimate is for each. The distinction is not cosmetic — a total conflates a student who is uniformly mediocre with one who is strong on nine objectives and absent on three, and those two students need entirely different lessons.

It carries uncertainty explicitly. One correct answer is weak evidence; five across three weeks is strong evidence. A model that distinguishes "probably knows this" from "has demonstrated this repeatedly" can direct assessment at what it does not yet know, which is how adaptive delivery decides what to ask next.

It has structure between objectives. Understanding rates of change requires understanding functions. When the model knows the prerequisite graph, a failure on the later objective is not just a fact about that objective; it is evidence about the earlier one, and often the earlier one is where the real gap is.

That last property is what makes prediction possible at all. Without a prerequisite structure, the best a model can do is describe the past.

Where the evidence comes from

The quality of any such model is bounded by the quality and frequency of its inputs, which is why this has historically been theoretical.

Termly summative assessment gives a handful of coarse data points a year — enough to describe a trajectory in retrospect, not enough to act on. What changes the picture is frequent, low-stakes, objective-mapped assessment: short checks that generate evidence continuously as a side effect of ordinary teaching, rather than an annual instrument that interrupts it.

Practice is the richer signal and the more delicate one. A student working through self-paced material produces far more evidence than any test — attempts, second attempts, which distractor they chose, what they abandoned. It is tempting to feed all of it into the model.

It should not be fed in as performance. The moment practice attempts influence anything a student is judged on, practice stops being practice: learners avoid what they might get wrong, and the most informative signal in the system — honest failure — disappears. Practice can inform what a student is offered next while remaining invisible as a record. That is a design constraint, not a preference, and a system that quietly relaxes it will produce a model that looks better and is worth less.

What it is good for

The honest applications are narrower than the phrase "digital twin" suggests, and still valuable.

Sequencing. If the model says a cohort is thin on a prerequisite that next week's topic depends on, next week's lesson can address it first. This is what experienced teachers do from intuition; the model makes it visible for thirty students at once, including the quiet ones whose gaps intuition misses.

Targeting scarce attention. Support time is limited and currently allocated by whoever is most visibly struggling. A model that flags a student whose estimate is deteriorating quietly — still passing, but drifting — surfaces the ones who do not present as a problem until they are one.

Distinguishing kinds of failure. Not knowing a fact, not recognising when a rule applies, and knowing the material but reading the question wrongly all look identical in a total score, and require different responses. Objective-level evidence separates them.

The claims worth refusing

Three overreaches recur in this area and are worth naming, because they are what make people rightly suspicious.

That the model knows the student. It knows a projection of the student onto a set of objectives somebody chose. Everything not in that set is invisible to it, including most of what makes a learner interesting.

That prediction justifies allocation. A model that predicts a student will struggle with a topic is a reason to teach it more carefully. It is not a reason to route them onto a lower pathway. Predictive sorting is how a support tool becomes a machine for confirming its own forecasts, and the effect is strongest on exactly the students who are already sorted against.

That more data is always better. Evidence collected from surveillance changes the behaviour it observes. A model built on watched practice is a model of what students are willing to be seen doing.

Where this sits today

The modelling is not the hard part. The prerequisite structures, the uncertainty tracking, the estimate updating are well-understood and have been for years.

The hard part has always been the evidence: enough assessment, mapped to objectives, marked to a consistent standard, generated often enough to keep a model current, without consuming every hour a teacher has. That is the constraint that has moved, and it moved because item generation and rubric-based scoring made frequent objective-mapped assessment affordable rather than aspirational.

Which means the interesting question is no longer whether a learner can be modelled usefully. It is what an institution is prepared to do with the model, and what it is prepared to refuse to do with it. That is a governance decision, and it should be made before the model is built rather than after it is persuasive.

AptitudAI — assessment that measures the learner, not the room.