All posts

Personalised learning app or just a recommendation engine? Six tests you can run yourself

2026-09-24

By The Sophoz Curriculum Team

The short answer. Most apps marketed as personalised only change the order of content. A genuinely adaptive system changes the difficulty and diagnoses the reasoning. You can tell which one you are looking at in about twenty minutes of a free trial, using six tests that require no technical knowledge. The quickest of them: get three questions deliberately wrong in a row and watch what arrives next. If it is three more questions of the same difficulty, you have a question bank with a scheduler.

Every product in this category uses the same vocabulary. Personalised, adaptive, learning path, AI-driven. The words do not distinguish anything, so this is a guide to testing rather than to reading the marketing.

The three things "personalised" usually means

Before the tests, it helps to know what you are separating.

Level 1: cosmetic personalisation. The app uses your child's name, remembers they like cricket, and puts cricket into the word problems. The mathematics is identical for every child. This is decoration, and it is not worthless, because a child who engages is a child who practises. It is just not what the word implies.

Level 2: content sequencing, which is a recommendation engine. The app decides what to serve next based on what your child has completed and what similar children did next. This is the mechanism behind streaming services and shopping sites, applied to lessons. It changes the order. The individual pieces of content are the same pieces every child gets, and the difficulty is not being adjusted to your child's current state, just the queue.

Most of the market lives here.

Level 3: adaptive difficulty and diagnosis. The app changes how hard the questions are while your child is working, and when your child gets something wrong it works out which underlying skill is missing and practises that instead.

The gap between Level 2 and Level 3 is the one that costs money and is never stated on a pricing page.


The six tests

Run these during a free trial, with your child's account or a test account. None takes more than five minutes.

Test 1: the three deliberate wrongs

What to do. Answer three questions in a row incorrectly, on purpose, on the same topic. Choose plainly wrong answers rather than near-misses.

What a recommendation engine does. Serves a fourth question of the same difficulty on the same topic. Possibly suggests a video on that topic. The queue is unchanged in character.

What an adaptive system does. Drops the difficulty, or moves to a prerequisite skill. If the topic was equivalent fractions, a good system will start asking about multiplication as scaling, because that is what equivalent fractions rest on.

Why this is the fastest test. Serving more of the same after repeated failure is the single clearest signal that nothing is being diagnosed. If a child does not understand something, twenty more instances of it is not a strategy.

Test 2: the three deliberate rights, answered fast

What to do. Answer three questions correctly and quickly.

What a recommendation engine does. Continues down the list. Your child completes the module, gets a tick, and moves to the next module.

What an adaptive system does. Skips ahead, or raises difficulty, within the session. A child demonstrating fluency should not be made to do seventeen more of the same.

What this catches. Systems that only adapt downwards. Plenty adjust when a child struggles and never adjust when a child is bored, which wastes the time of exactly the children most likely to benefit.

Test 3: the wrong-answer explanation

What to do. Get a question wrong and read what the app says about it.

What a recommendation engine does. Shows the correct answer, and possibly the correct working. "The answer is 3/4."

What an adaptive system does. Addresses why that particular wrong answer was chosen. If your child answered 7/12 when adding 1/3 and 1/4, the useful response names the misconception: the numerators and denominators were added separately, which treats a fraction as two numbers rather than one.

Why this matters more than it looks. A wrong answer is a window into a method. A system that only shows the right answer has closed the window. This is the difference between marking and teaching, and it is the single most informative test on this list.

Test 4: the hint test

What to do. Ask for a hint on a question you could solve, and see what arrives. Then ask again.

What a recommendation engine does. Gives the answer, or gives the first step of the answer, which usually amounts to the same thing.

What an adaptive system does. Gives a hint that says what to try next and why that is a sensible thing to try, without doing it. Research on tutoring systems distinguishes principle-based hints, which state what to do next and why in terms of the underlying problem-solving principles, from bottom-out hints, which give the answer and turn the question into a worked example.

A nuance worth knowing. Bottom-out hints are not automatically bad. One study found time spent on bottom-out hints correlated positively with learning, because some students deliberately reach the answer and then explain it to themselves. The failure is not the hint, it is not thinking about it. What you are testing for is whether the system has more than one kind of hint.

A sobering statistic to keep expectations honest. In research on tutoring systems, students viewed 68 % of non-final hint levels for less than one second, too fast to have read them. Whatever hint design a product has, most of it will be clicked past.

Test 5: the parent report test

What to do. Ask the company to show you a real parent report. Not a screenshot from the marketing site; a real one, from a real week.

What a recommendation engine produces. Percentages, progress bars, time on task, topics completed, a streak. All of it is counting.

What an adaptive system produces. A sentence a parent can act on. Something closer to: "she is converting improper fractions correctly but divides instead of multiplying when scaling a recipe up, which suggests the scaling idea is not yet secure."

Why this is the test companies find hardest to fake. A system can only report a misconception if it has identified one. A report full of percentages is not a design choice about parent preference. It is usually the ceiling of what the system knows.

Test 6: the mixing test

What to do. Look at a practice session and ask: are all these questions on the same topic?

What most apps do. Follow the chapter. Twenty questions on ratio, then twenty on averages, because that is how the syllabus is organised and how content was authored.

What a well-designed system does. Mixes problem types once a method is secure, so the child has to work out what kind of question this is before working out the answer.

Why this is worth checking even though it sounds like a detail. It is the strongest finding in the practice literature. In a randomised controlled trial with 787 seventh-graders across 54 classes over four months, interleaved practice produced 61 % on a delayed test against 38 % for blocked practice, an effect size of d = 0.83 (Rohrer, Dedrick, Hartwig and Cheung, 2020). Chapter-ordered practice is blocked practice by construction, whatever else the app does well.


The scorecard

Test Recommendation engine Adaptive system
1. Three deliberate wrongs Another similar question Drops to a prerequisite skill
2. Three fast rights Continues the list Raises difficulty mid-session
3. Wrong-answer explanation Shows the correct answer Names why that wrong answer was chosen
4. Hint quality Gives the answer or first step Says what to try and why, without doing it
5. Parent report Percentages and streaks A named misconception in a sentence
6. Practice structure Follows the chapter Mixes types once a method is secure

Four or more in the right-hand column means the system is doing something real.

Two or three is the common case: partly adaptive, usually on difficulty but not on diagnosis. Whether that is worth paying for depends on what you need.

Zero or one means you are buying a well-organised question bank, which is a legitimate product. Just compare its price against a question bank, not against a tutor. Free options exist, and a set of workbooks with good pattern fit runs to a few hundred rupees.

Three claims to discount entirely

"AI-driven learning path." Every product that queues content can say this. It describes the existence of a queue.

"Personalised for your child." Tests 1 and 3 settle this in four minutes. Do not settle it by reading.

"Proven to improve results by X %." Ask the obvious follow-up: compared with what, over how long, measured how, and who counted. Every improvement figure in this category, including ours, is self-reported and unaudited. We could not find a single product in the Indian market publishing an independently audited outcome comparison. That does not make the numbers false. It means they are marketing rather than evidence, and they should be weighted accordingly.

What no app can do, at any level

Worth saying plainly, because it is the boundary of the whole category.

The system sees answers. It does not see that your child has decided they are bad at maths, that the last fortnight has been hard, or that they are copying a friend's answers to get through a session faster. Those are usually the things most worth knowing.

There is also a limit that shows up in the research. A tutoring system explicitly designed to teach children better help-seeking produced durable behavioural improvement that persisted for months, and no improvement in what they actually learned (Aleven, Roll, McLaren and Koedinger, 2016). Software can change behaviour inside the software more easily than it can change learning.

So the useful question is not which app is most adaptive. It is what specific problem you are trying to solve. If the problem is that nobody knows why the same error keeps recurring, adaptive diagnosis is the right tool and the tests above will find one. If the problem is motivation, or confidence, or that nobody has sat beside the child in a month, no amount of adaptation addresses it.

Frequently asked questions

What is the difference between personalised and adaptive learning? Personalised usually means changing the order or presentation of content, such as recommending the next video or theming word problems around a child's interests. Adaptive means changing the difficulty and diagnosing which underlying skill is missing, based on the child's answers. Most products described as personalised are doing the first.

How can I tell if a learning app is genuinely adaptive? Get three questions deliberately wrong in a row. A question bank serves a fourth similar question; an adaptive system drops to a prerequisite skill. Then get one wrong and read the explanation: showing the correct answer is marking, naming why that particular wrong answer was chosen is diagnosis.

Are personalised learning apps worth paying for? It depends which level you are buying. A well-organised question bank is a legitimate product but should be priced against question banks, not against tutoring. Genuine adaptive diagnosis solves a specific problem, which is identifying the missing prerequisite behind a recurring error, and is worth paying for when that is the problem you have.

What should a good parent report from a learning app contain? A sentence naming a specific misconception, ideally quoting what the child actually wrote. Percentages, progress bars, streaks and time on task are counting rather than diagnosis, and a report limited to them usually indicates the limit of what the system knows.

Do learning apps improve exam results? The best evidence for the category is VanLehn's 2011 review, which found step-based intelligent tutoring systems produced roughly the same effect as human tutoring, d = 0.76 against d = 0.79. Individual products' own improvement claims are self-reported and unaudited, and should be treated as marketing until someone publishes a controlled comparison.

Does it matter if an app follows the chapter order? Yes, more than most parents expect. Chapter-ordered practice is blocked practice, and a randomised trial with 787 seventh-graders found interleaved practice produced 61 % on a delayed test against 38 % for blocked. An app that only walks the syllabus in order has automated a workbook.


Sophoz builds adaptive practice and would be one of the products these tests are applied to, which is a reason to apply them to us as well. Research citations are linked in full in what the research says about olympiad preparation.