What the research actually says about preparing for a maths olympiad
By The Sophoz Curriculum Team
The short answer. Four things have good evidence behind them: practise retrieving rather than rereading, space sessions out rather than massing them, mix problem types rather than drilling one at a time, and study worked solutions early in a new topic. Two things commonly recommended have weaker evidence than you have been told: letting young children struggle before instruction, and counting hours of practice. And one popular question has no researched answer at all, which is worth knowing before someone sells you one.
This is a longer piece than most, because the useful part is not the list of techniques. It is the caveats, and where the caveats mean the standard advice is wrong for a nine-year-old specifically.
Mixing problem types is the single strongest finding, and it makes practice look worse
If you change one thing about how your child practises, change this.
In a randomised controlled trial with 787 seventh-graders across 54 classes in a large Florida district, running over four months under normal classroom conditions with no researcher training of teachers, students whose practice problems were interleaved scored 61 % on an unannounced test a month later. Students whose practice was blocked by topic scored 38 %. The effect size was d = 0.83 (Rohrer, Dedrick, Hartwig & Cheung, 2020, Journal of Educational Psychology, full text).
That is a large effect under realistic conditions, and the study is independently listed by the US Institute of Education Sciences What Works Clearinghouse. It replicates earlier findings: 126 seventh-graders scored 74 % interleaved against 42 % blocked on a thirty-day test (Rohrer, Dedrick & Stershic, 2015), and 140 seventh-graders scored 72 % against 38 % on a two-week test with d = 1.05 (Rohrer, Dedrick & Burgess, 2014). That last pair of figures is very widely misattributed to the 2015 paper; it is the 2014 one.
Blocked practice is what a workbook does by default: twenty questions on ratios, then twenty on averages. Interleaved practice shuffles them, so the child has to work out what kind of problem this is before working out the answer.
Here is why it matters for olympiad preparation specifically, and it is the most useful single sentence in this literature. In Rohrer and Taylor's 2007 study, blocked practice produced 89 % accuracy during the practice session and 20 % a week later. Mixed practice produced 60 % during practice and 63 % a week later (Rohrer, 2009, JRME).
Read that twice. The method that looked far better while the child was doing it was the one that had almost entirely evaporated a week later. Practice-session accuracy and durable learning pointed in opposite directions.
This has a direct consequence for how parents judge preparation. A child working through a blocked worksheet gets most of them right, feels competent, and the parent sees a session going well. A child working through mixed problems gets more wrong, hesitates more, and looks like they are struggling. On this evidence the second child is learning more, and the visible signal is misleading in a way that predictably pushes families towards the worse method.
Why it works. Blocked practice lets a child skip the hardest step. When every problem on the page is a ratio problem, they never have to decide that it is a ratio problem. Mixing forces that decision every time. And deciding what kind of problem this is, under time pressure, on a paper where the next question could be anything, is precisely what a SOF paper asks, and it is the entire skill an olympiad paper tests.
The caveat. These studies are seventh-graders and specific topics, and the authors state that generalisability across a wider variety of material, teachers and students is unknown. The most age-relevant study, on fourth-graders, reports d = 1.23 but we could not retrieve it directly, so treat that figure as secondhand.
Space sessions out, but the popular rule is more precise than the evidence
Across 317 experiments and 14,811 participants, spacing study sessions rather than massing them improved final performance in 259 of 271 comparisons (Cepeda, Pashler, Vul, Wixted & Rohrer, 2006, Psychological Bulletin, full text).
For scheduling, a later study of over 1,350 participants, with gaps up to 3.5 months and tests up to a year later, found the optimal gap was roughly 20 % of the test delay for delays of a few weeks, falling to roughly 5 % for a one-year delay (Cepeda, Vul, Rohrer, Wixted & Pashler, 2008). The authors' conclusion is worth quoting: "The interaction of gap and test delay implies that many educational practices are likely to be highly inefficient."
What that means in practice. For an exam ten weeks out, revisit each topic roughly every two weeks. For an exam a year out, roughly every three weeks. Which is to say: the gap grows with the horizon, but nothing like proportionally.
Where the popular version overstates it. The rule of thumb circulating as "review at 10 to 20 % of your retention interval" flattens a ratio that actually declines as the horizon lengthens. It is derived from verbal recall and paired-associate learning, not from mathematical problem solving, and the retention function is fairly flat near its optimum, so exactness buys you nothing. The 2006 meta-analysis is also weaker support for long-horizon planning than it is usually made to carry: 80 % of the performance differences it aggregates used a retention interval of under one day, only 4 % exceeded a month, and the authors explicitly note insufficient data on children's long-term retention.
Say "roughly every fortnight" and mean it. Do not buy a schedule built to the day.
Retrieval beats rereading, and the gap widens with time
Students who read a passage then tested themselves on it recalled 56 % a week later. Students who read it four times recalled 42 %. At five minutes the ordering reversed: the repeated readers were ahead, 81 % to 75 % (Roediger & Karpicke, 2006, Psychological Science, full text).
And the repeated-study students predicted they would remember more, and remembered less.
The largest review of study techniques rated ten common methods for utility. Only two earned a high rating: practice testing and distributed practice. Rereading and highlighting, the two things students actually do most, both landed in the low-utility tier, alongside summarisation and keyword mnemonics (Dunlosky, Rawson, Marsh, Nathan & Willingham, 2013, Psychological Science in the Public Interest).
| Technique | Utility rating |
|---|---|
| Practice testing | High |
| Distributed practice | High |
| Elaborative interrogation | Moderate |
| Self-explanation | Moderate |
| Interleaved practice | Moderate |
| Summarisation | Low |
| Highlighting and underlining | Low |
| Keyword mnemonic | Low |
| Imagery for text learning | Low |
| Rereading | Low |
Two caveats that matter. "Low utility" in that paper often means insufficiently evidenced or narrowly applicable rather than proven useless. And interleaving was rated only moderate in 2013 because the evidence base was thin at the time; the strong mathematics trials described above mostly postdate it. Do not cite Dunlosky against interleaving.
The applied point for olympiad work: reading a worked solution and nodding is recognition, not retrieval. Covering the solution and reconstructing it is retrieval. The second takes four times as long and feels worse.
Worked examples work, and they stop working exactly where olympiads begin
This finding comes with a limitation built into the original study, and the limitation is the part that matters here.
Students who alternated between studying worked algebra solutions and solving problems spent roughly a sixth as long in acquisition, 32 seconds per problem against 185 seconds, and afterwards solved test problems faster and with fewer errors than students who solved everything themselves. With acquisition time held equal, the worked-example group got through about 25 problems against 8 (Sweller & Cooper, 1985, Cognition and Instruction, full text).
Then Experiment 4 in the same paper: the worked-example advantage appeared on similar test problems and vanished on dissimilar ones. Sweller and Cooper attribute this to the specificity of the schemas being acquired.
Worked examples buy fluency on the problem type studied. They do not by themselves buy transfer. And transfer to unfamiliar problem structures is the entire content of an olympiad paper. So worked examples are the right tool at the start of a new topic and the wrong tool as the only tool.
Two refinements make this usable:
Expertise reversal. Instructional support that helps a novice becomes redundant and then actively harmful as the learner gains expertise, because processing the redundant guidance itself costs attention (Kalyuga, Ayres, Chandler & Sweller, 2003). The same worked-example-heavy method that is right for a child starting a topic is wrong for that child eighteen months later. There is no correct ratio of examples to problems; the ratio has to move.
Fading. The operational answer is to remove solution steps one at a time, so a full worked example becomes a completion problem and then an unaided problem (Renkl, Atkinson & Große, 2004). The mechanism is well established; we did not verify effect sizes for it.
The productive struggle advice is wrong for younger children, specifically
This is the finding most likely to contradict what you have been told, and the age boundary is the point.
Productive failure, where students attempt a problem before being taught the method, has real support. Ninth-graders who spent two lessons inventing their own ways to measure variance before being taught the formula were no worse at procedure and substantially better at conceptual insight, 16.40 against 8.20, partial η² = .27 (Kapur, 2012, Instructional Science).
The meta-analysis covering 45 articles, 53 studies and 166 comparisons finds Hedges' g = 0.36 for conceptual knowledge and transfer, and a null result for procedural knowledge, g = −0.03 (Sinha & Kapur, 2021, Review of Educational Research, full text).
And then the finding that almost never reaches parents: effects were negative for students in grades 2 to 5. For those children, instruction-first outperformed problem-solving-first. Positive effects increased with grade level. The authors suggest younger learners may lack the metacognitive strategies productive failure requires, and that scaffolding the initial problem-solving phase might be critical.
So for a child in class 3, 4 or 5, the evidence points towards teaching the method and then practising it, not towards handing them an unfamiliar problem and waiting. Productive failure is not a universal prescription for children, and below roughly age eleven the evidence points the other way.
It is also highly fidelity-dependent. The four strongest predictors of it working were instruction building on students' own solutions, group-work structure, evidence of multiple solution attempts, and dialogue-dominant facilitation. Simply putting a hard problem before instruction is not productive failure; it is just a hard problem.
Nobody knows how long a child should struggle before getting a hint
This is the question parents ask most and the one where confident answers are least warranted.
The researchers who have studied it most closely named the problem and left it open. Koedinger and Aleven call it the assistance dilemma: "How should learning environments balance information or assistance giving and withholding to achieve optimal student learning? How best to achieve this balance remains a fundamental open problem in instructional science" (2007, Educational Psychology Review).
Any article telling you to let a child struggle for fifteen minutes is asserting a number the field's own leading researchers describe as unresolved.
What the research on hints does establish is more interesting than a number:
- Students viewed 68 % of non-final hint levels for less than one second, too fast to have read them.
- Conversely, even after three errors on a step, students asked for a hint only 34 % of the time. Help avoidance is as common as help abuse.
- A tutoring system explicitly designed to teach better help-seeking produced durable improvement in help-seeking behaviour, persisting months after feedback ended, and no improvement in what students actually learned (Aleven, Roll, McLaren & Koedinger, 2016, full text).
The distinction that does hold is between kinds of hint. A principle-based hint says what to do next and why that is a sensible thing to do. A bottom-out hint gives the answer, which converts the problem into a worked example. Both have a place. Shih, Koedinger and Scheines found time spent on bottom-out hints correlated positively with learning, because some students deliberately reach the answer and then explain it to themselves. A bottom-out hint is not automatically a failure. Not thinking about it is the failure.
The practical version: the question is not how long to wait, it is whether the child is still generating attempts. A child trying a smaller case, drawing something, or testing a guess is working. A child staring is not, and waiting longer does not change that.
The heuristics are worth teaching, but "research shows Polya works" is not a claim you can make
Polya's four phases, understand the problem, devise a plan, carry it out, look back, and the heuristics that go with them, solve a smaller case, draw a figure, work backwards, look for a pattern, consider a related problem, are the vocabulary of competition mathematics for good reason.
But How to Solve It (1945) is a normative work by a mathematician. It has no sample, no control group and no effect size. Citing Polya is citing an argument, not evidence.
Schoenfeld took up the empirical question and his answer is more complicated. His claim is that success depends on four things, knowledge, heuristics, control, meaning the metacognitive decisions about what to pursue and when to abandon it, and belief systems, and that control is what most distinguishes competent from incompetent solvers. His critique of Polya is that the heuristics as stated are descriptively accurate but too coarse to execute: "consider a related problem" names a family of much finer strategies that each have to be taught separately (Schoenfeld, 1992, full text).
Note that the widely-quoted percentages from Schoenfeld's protocol analyses, about how long novices spend in an undirected exploration phase, we were unable to verify from source and therefore do not reproduce.
Two numbers you should stop repeating
The 10,000-hour rule is not a research finding. Ericsson, Krampe and Tesch-Römer (1993) reported that the best violinists had accumulated 7,410 hours by age 18, against 5,301 and 3,420 for two comparison groups. That was one group mean, not a threshold. The 10,000-hour framing comes from Gladwell's Outliers, and Ericsson objected to it repeatedly.
More significantly: the founding study's central finding did not replicate. In a 2019 replication with 39 violinists, the best violinists had accumulated less practice than the merely good ones, 8,224 hours against 9,844. Deliberate practice explained 26 % of variance among experts in the replication against 48 % in the original (Macnamara & Maitra, Royal Society Open Science).
And the meta-analysis: across music, games, sports, education and the professions, deliberate practice explains a minority of performance variance, and in education it explains about 4 to 5 % (Macnamara, Hambrick & Oswald, 2014, Psychological Science). The moderators strengthen the sceptical reading rather than weakening it. Studies using retrospective questionnaires found 15 %; studies using ongoing diary logs, the more valid measure, found 4 to 5 %. Highly predictable activities showed 23 %; low-predictability activities showed 6 %. Competition mathematics, where novelty is the entire point, sits at the unpredictable end.
The honest position: structured practice is necessary and it matters. It is not sufficient, its measured contribution shrinks as the measurement improves, and "enough hours will get any child to a medal" is not something the evidence supports.
One-to-one tutoring is worth about 0.8 standard deviations, not 2. Bloom's famous two-sigma figure does not survive scrutiny. VanLehn's review puts human tutoring at d = 0.79, step-based intelligent tutoring systems at d = 0.76, and the difference between them at d = 0.21, not reliably different (VanLehn, 2011, Educational Psychologist). Bloom's tutors were minimally-trained undergraduates, and the tutored condition used a 90 % mastery threshold against 80 % for the comparison, which inflated the gap. Still one of the largest effects in education. Less than half the number that gets quoted.
What this adds up to, as a week
Nothing here requires a platform. It requires a shape.
- Start a new topic with worked solutions, not problems. Read the solution, cover it, reconstruct it. Two or three examples, then fade: leave out the last step, then the last two.
- Mix as soon as the method is secure. Once the child can do five ratio problems in a row, stop doing ratio problems in a row. Twelve mixed problems beats thirty blocked ones, and it will look worse while they do it.
- Revisit old topics roughly every fortnight for an exam a couple of months out. Not because a schedule says so, but because everything else in this piece depends on the material still being there.
- Test, do not review. A ten-minute attempt from a blank page is worth more than forty minutes of reading through what was done last week.
- For a child under about eleven, teach the method first. The productive-failure evidence turns negative for grades 2 to 5.
- Spend a quarter of the time on the Achievers Section, because it is a quarter of the marks in five questions. The arithmetic is in our SOF exam guide.
- Judge hints by what happens after them. Whether the child then explains the step back, unprompted, matters more than how long they held out.
This is also, roughly, the shape Sophoz's olympiad practice is built to: worked examples with a visual walk-through for new material, mixed practice rather than topic-blocked drills, a schedule that works backwards from the exam date and brings earlier topics round again, and an assistant that answers with a hint or another way of seeing it rather than the answer. Sophoz publishes this research because the method is not proprietary. It predates every platform selling it, including ours, and a family with a workbook and a calendar can run it.
Frequently asked questions
What is the single most effective way to practise for a maths olympiad? Mixing problem types rather than drilling one type at a time. In a randomised trial with 787 seventh-graders, interleaved practice produced 61 % on a delayed test against 38 % for blocked practice, d = 0.83. It will make practice sessions look worse while producing better results a month later.
How long should my child struggle with a hard problem before getting help? There is no researched answer. Koedinger and Aleven, who have studied this most closely, call the balance between giving and withholding assistance "a fundamental open problem in instructional science". A better question than how long is whether the child is still generating attempts.
Should my eight-year-old work things out before being taught? On current evidence, usually not. The meta-analysis of productive failure finds effects that are positive overall but negative for grades 2 to 5, where instruction-first outperformed problem-solving-first. The advice to let children discover methods themselves appears to be wrong for younger children specifically.
How many hours a day does olympiad preparation take? Nobody has established a number, and the hours framing is weaker than it sounds. In education, accumulated deliberate practice explains roughly 4 to 5 % of performance variance, and the better the measurement of practice, the smaller the estimate becomes. Shape of practice matters more than quantity: forty mixed minutes beats ninety blocked ones.
Is rereading notes useless? Not useless, but among the weakest of the common techniques. Dunlosky et al. rated rereading and highlighting low utility while rating practice testing and distributed practice high. In Roediger and Karpicke's study, repeated reading beat self-testing at five minutes and lost by fourteen percentage points at one week.
Do worked examples help or make children dependent? They help substantially at the start of a topic and their advantage disappears on problems unlike the ones studied. Sweller and Cooper's own fourth experiment found the benefit vanished on dissimilar problems. Fade them out as the child gains ground.
Every study cited here is linked. Where we could not verify a figure from the source, we have said so rather than reproduce it.