Stop doing 30 time-speed-distance questions in a row. In a pre-registered randomised controlled trial of 787 students across 54 maths classes, the group whose assignments mixed problem types scored 61% against 38% for the group who practised one topic at a time, on an unannounced test a month later - an effect size of d = 0.83. In my own 20-mock log, the section I practised topic-wise (QUANT) rose from a 26.8 average across my first five mocks to 35.6 across my last five, while DILR - the section where every question is a selection decision - moved from 18.4 to 20.6. Two marks in three months.
Those two facts point in the same direction, and it is not the one you would guess. Blocked practice builds execution. Interleaved practice builds selection. CAT scores selection.
The trial
Doug Rohrer, Robert Dedrick, Marissa Hartwig and Chi-Ngai Cheung, University of South Florida, published in the Journal of Educational Psychology in 2019. It is a pre-registered, cluster randomised controlled trial - the design you want and almost never get in education research.
Fifty-four seventh-grade mathematics classes periodically completed either interleaved or blocked assignments over four months. In a blocked assignment every problem uses the same strategy. In an interleaved assignment, no two consecutive problems can be solved the same way. Both groups then completed an identical interleaved review assignment, and a month after that, students sat an unannounced test.
| Study | Who, and how many | Outcome measured | Interleaved | Blocked | Effect |
|---|---|---|---|---|---|
| Rohrer et al. 2019 | 787 seventh-graders, 54 classes, 4 months | Unannounced test, 1 month later | 61% | 38% | d = 0.83 |
| Samani & Pan 2021 | 350 undergraduates, physics, 8 weeks | Surprise test, stage 1 | 0.54 | 0.43 | d = 0.40 |
| Samani & Pan 2021 | Same students, stage 2 | Surprise test, stage 2 | 0.47 | 0.27 | d = 0.91 |
| Samani & Pan 2021 | Same students | Their actual course midterms | - | - | d = 0.02, p = .88 |
| Foster et al. 2019 | 126 undergraduates, geometry volumes | Final test, all problems | - | - | d = 0.67 |
Take the last row of Samani and Pan seriously. It is in the table because leaving it out would be the standard move, and it is the most useful line in the whole post. More on it below.
Why interleaving works, and why that maps onto CAT
Rohrer et al. explain the size of their effect by pointing out that interleaved practice is really three techniques at once. Their reasoning, in their words:
"First, the mixture of different kinds of problems within each assignment provides students with an opportunity to practice choosing a strategy on the basis of the problem itself, which is precisely what students must do when they encounter a problem on a cumulative exam or other high-stakes test."
The other two are that interleaving forces spacing of each individual skill across assignments, and that it pushes students toward recalling a method rather than reading it off the top of the page.
That first mechanism is the one that should stop a CAT aspirant. A blocked practice set tells you what kind of problem it is before you read it. The chapter heading answers the hardest question in the paper. On the real paper, nothing does - and choosing what to solve, in what order, and what to abandon is most of what separates a 60 from a 90 in a section like DILR.
Rohrer et al. make the same observation about their own materials: a problem may contain no obvious cue such as "hypotenuse" or "right triangle", so a student who has only ever practised in blocks has never had to infer which tool applies. Interleaved practice, in their phrase, "requires students to choose a strategy and not merely execute a strategy".
My own data does not prove this, and that is why it is interesting
Comparing my first five CAT mocks against my last five, over roughly three months:
| Section | First 5 mocks | Last 5 mocks | Change | How I practised it |
|---|---|---|---|---|
| QUANT | 26.8 | 35.6 | +8.8 | Topic-wise problem sets |
| VARC | 16.6 | 24.8 | +8.2 | Passages, largely unstructured |
| DILR | 18.4 | 20.6 | +2.2 | Sets, solved one after another |
Read naively, this is evidence against interleaving: the section I drilled most blockedly improved most. I am not going to pretend otherwise.
The honest reading is different. QUANT rewards execution, and blocked practice is an efficient way to build execution - which is exactly what Rohrer et al.'s third caveat says, that some blocked practice on a new skill is probably necessary before interleaving is useful. What my quant improvement did not buy me was selection, and the place that showed was DILR.
On 29 July I scored 96 overall, my best paper to that point. On 30 July I scored 40, with 2 marks in DILR. My own analysis note from that week reads: "Could have definitely gotten 30 marks." The sets were doable. I picked the wrong one, committed early, and ran out of clock. Nothing about my knowledge changed in 48 hours; what varied was triage, and I had never practised triage, because every practice session I ran told me in advance what I was solving. The whole log is in the 20-mock post.
How to interleave CAT quant without wrecking your prep
Rohrer et al. list four caveats on their own result. All four are practical instructions, so I am using them as the structure here rather than inventing advice.
- Learn a topic blocked, then mix it in. Caveat three: interleaved practice "might be less effective or too difficult if students do not first receive at least a small amount of blocked practice when they encounter a new skill or concept". Their interleaved group had already had blocked instruction from their teachers. Do not interleave a topic you have not yet learned.
- Expect it to take longer per question. Caveat one: every teacher in the trial reported the interleaved assignments took more time, and the authors note the effect "would have been smaller if it had been measured per unit of time invested". Nobody has measured interleaving's benefit per hour. A mixed set of 20 is not a substitute for a blocked set of 20 in the same slot.
- Check your answers and correct them. Caveat four: interleaved practice "might be effective only if students receive corrective feedback". Every study in this literature provided it. An unmarked mixed set is not the intervention that was tested.
- Do not judge it by tomorrow's accuracy. Caveat two: the benefit "might be smaller at shorter test delays", and the one prior study that found no interleaving effect used a test delay of 2-5 days. Interleaved practice makes you look worse on the day and better a month later, which is a hard trade to hold your nerve on.
Concretely, for CAT: keep learning new topics in blocks, then build a weekly mixed set of 20-25 questions drawn from every quant topic you have covered so far, in random order, timed, and mark it. Sectional mocks do some of this for you, which is one reason a weekly sectional beats an extra topic worksheet once your syllabus is broadly covered.
The result that complicates all of this
Joshua Samani and Steven Pan ran interleaving in a real undergraduate physics course at UCLA over eight weeks, with 350 students across two lecture sections completing thrice-weekly homework. On surprise criterial tests the interleaved group's median scores were 50% and 125% higher, at d = 0.40 and d = 0.91.
On the students' actual high-stakes midterms, three days after each criterial test, there was no significant difference at all - d = 0.20 (p = .09) after stage one and d = 0.02 (p = .88) after stage two.
The authors offer a reasonable explanation: exit surveys confirmed most students crammed heavily before the midterms and not before the surprise tests, and mean midterm accuracy was 0.74 against 0.42 on the criterial tests, so the exams may simply have been too easy and too crammed-for to show the difference. They also say plainly that if such benefits do not reliably appear on high-stakes exams, "then that would constitute a notable limitation, particularly if enhancing exam performance was the sole objective."
For a CAT aspirant, enhancing exam performance is the sole objective. So the strongest classroom evidence for interleaving comes with an unresolved question about whether it shows up on the exam that counts. I have not seen another exam-prep article mention this, and it is in the abstract's own paper.
And the result that suggests interleaving is partly just spacing
Nathaniel Foster and colleagues at Kent State ran 126 undergraduates through three conditions on geometry volume problems: blocked, interleaved, and remote interleaved - which alternated volume problems with completely unrelated ones like permutations and fraction addition.
If interleaving works because mixing similar problem types lets you notice the contrasts between them, remote interleaving should not help. It helped more. Standard interleaving beat blocked practice on the final test at d = 0.67; remote interleaving beat standard interleaving at d = 1.01 and blocked at d = 2.19. The authors report that "standard interleaving alone did not beat remote interleaving on any of the dependent measures".
Their conclusion is that a substantial part of the interleaving benefit is simply distributed practice - spacing each problem type out - rather than the discriminative contrast most explanations lead with. Practically, that is reassuring for CAT: it means a mixed set does not need to be cleverly composed of confusable topics. Mixing at all is most of the point.
Where this evidence does not reach
- 787 seventh-graders are not CAT aspirants. Twelve and thirteen year olds, learning school algebra, on four-month timescales, in a US classroom. The mechanism - having to choose a strategy from the problem itself - transfers cleanly as a description. The 61-to-38 gap does not transfer as a prediction about your mock score.
- Nobody has measured the benefit per hour. Rohrer et al. say so explicitly. Interleaved practice is slower per question, and every serious CAT aspirant is time-constrained, so the practically relevant effect size is one that has never been calculated.
- The high-stakes-exam evidence is genuinely weak. The UCLA null on midterms is the only direct test I could find of whether classroom interleaving shows up on a real exam, and it did not, for reasons the authors can only speculate about.
- The mechanism is contested. Foster et al.'s dissociation suggests interleaving may work mostly through spacing. That does not undermine the advice, but it should make you suspicious of anyone explaining the mechanism confidently.
- My log is one person and cannot separate the variables. QUANT improved most and was practised in blocks. DILR barely moved and was also practised in blocks. Both facts are consistent with several stories, and I am choosing the one the literature supports rather than one my data proves.
- Nothing here was tested on DILR. The claim that set selection is a trainable skill responsive to mixed practice is my inference from the strategy-choice mechanism, not a finding. No study I have read tested anything resembling a CAT logical-reasoning set.
Sources
- Rohrer, D., Dedrick, R. F., Hartwig, M. K., & Cheung, C.-N. (2019). A randomized controlled trial of interleaved mathematics practice. Journal of Educational Psychology. Read in full via ERIC. All figures and all four caveats above are from that text.
- Samani, J., & Pan, S. C. (2021). Interleaved practice enhances memory and problem-solving ability in undergraduate physics. npj Science of Learning, 6:32. Open access.
- Foster, N. L., Mueller, M. L., Was, C., Rawson, K. A., & Dunlosky, J. (2019). Why does interleaving improve math learning? Memory & Cognition, 47, 1088-1101.
- Mock log, 20 attempts, 10 May to 8 August 2026 - first-party, published in full in the 20-mock post.
If you want to see whether mixed practice is doing anything for you, the number to watch is not accuracy on the practice set - it will fall - but your sectional score four to six weeks later. Karma Yogi keeps sectional scores by date for exactly that reason, and the spacing side of the same argument is a separate post.
End of essay
- Anish Guruvelli