Revise a mock more than once, and put days between the passes rather than hours. The best evidence for that is Cepeda, Pashler, Vul, Wixted and Rohrer's 2006 meta-analysis in Psychological Bulletin, which pooled 14,811 participants and found spaced study produced 47.3% correct on the final test against 36.7% for massed study. In my own log of 20 CAT mocks taken between 10 May and 8 August 2026, the median gap between two mocks was 6 days - but 7 of the 19 gaps were two days or less, and those clustered attempts are the ones that taught me nothing.
That is the honest headline. What follows is the rest of it, including the part where the same meta-analysis fails to support the expanding-interval schedule that every spaced-repetition app on the market - including the one I built - actually uses.
What the meta-analysis actually found
The paper is Cepeda et al. (2006), "Distributed practice in verbal recall tasks: A review and quantitative synthesis". Its own description of its scope: "This review found 839 assessments of distributed practice in 317 experiments located in 184 articles." It is the reason anybody in education says the words "spacing effect" with confidence.
Their Table 1 compares massed against spaced presentation, split by how long the gap was between studying and the final test - the retention interval. This is the table that matters, and almost nobody who cites the paper reproduces it, because the middle of it is inconvenient.
| Retention interval | Massed | Spaced | Studies | Participants | Significance |
|---|---|---|---|---|---|
| 1-59 seconds | 41.2% | 50.1% | 96 | 5,086 | p < .001 |
| 1 min to under 10 min | 33.8% | 44.8% | 117 | 6,762 | p < .001 |
| 10 min to under 1 day | 40.6% | 47.9% | 10 | 870 | p = .535 |
| 1 day | 32.9% | 43.0% | 15 | 1,123 | p = .249 |
| 2-7 days | 31.1% | 45.4% | 9 | 435 | p = .190 |
| 8-30 days | 32.8% | 62.2% | 6 | 492 | p < .05 |
| 31 days or more | 17% | 39% | 1 | 43 | - |
| All intervals | 36.7% | 47.3% | 254 | 14,811 | p < .001 |
Three things fall out of that table that you will not read anywhere else in Indian exam-prep writing.
1. The overall effect is large and extremely well evidenced. A 10.6 percentage-point gap across 254 studies and nearly 15,000 people is not a marginal finding. If you take nothing else from this post, take that spacing is real.
2. The effect is best evidenced at the two ends and thinnest in the middle. Look at the rows for 1 day and 2-7 days: p = .249 and p = .190. Those are not statistically significant. The raw gaps look big - 32.9 against 43.0, and 31.1 against 45.4 - but only 15 and 9 studies existed in each bin. The 8-30 day row, where spacing nearly doubles performance (32.8% to 62.2%), rests on six studies and 492 people.
3. The bins that carry the statistical weight are the ones irrelevant to you. More than 11,000 of those 14,811 participants sat in the two rows where the final test came less than ten minutes after studying. A CAT aspirant's retention interval is measured in months.
The one rule the paper does support strongly
Cepeda et al.'s central claim is not "space your revision by N days". It is that the optimal gap grows as the interval you need to remember over grows. Their own summary of the practical position is unusually blunt:
"After more than a century of research on spacing, much of it motivated by the obvious practical implications of the phenomenon, it is unfortunate that we cannot say with certainty how long the ISI should be to optimize long-term retention."
They go on: for most practical purposes the retention interval is months or years, "so the optimal ISI will likely be well in excess of one day."
For a CAT aspirant that translates into something usable. If you sit a mock in June and the exam is in late November, your retention interval is roughly five months. Nothing in this literature tells you the optimal gap for a five-month retention interval - the entire evidence base for retention intervals longer than a month is one study with 43 participants. What it does tell you is that a gap measured in hours is almost certainly too short, and that a gap measured in days is more defensible than one measured in minutes.
Where I have to argue against my own product
Karma Yogi runs revision rounds on an expanding schedule. R0 is the mock itself, then R1 three days later, R2 seven days after R1, R3 fourteen days after R2, and so on out to R7. Each interval runs from when you completed the previous round, not from the mock date.
| Round | Gap from previous round | Days after the mock, if you never slip |
|---|---|---|
| R0 | The paper itself | 0 |
| R1 | 3 days | 3 |
| R2 | 7 days | 10 |
| R3 | 14 days | 24 |
| R4 | 21 days | 45 |
| R5 | 30 days | 75 |
| R6 | 45 days | 120 |
| R7 | 60 days | 180 |
That expanding shape is the standard design across every spaced-repetition tool. Cepeda et al. do not support it over a fixed schedule. Their Table 8 compares expanding against fixed intervals across 18 studies and 1,518 participants: 62.0% for expanding against 58.6% for fixed, t(42) = 0.5, p = .61. Not significant. Their assessment of the wider literature is harsher still:
"Some researchers have suggested, with little apparent empirical backing, that expanding inter-study intervals improve long-term learning... Our review of the evidence suggests that, in general, expanding intervals either benefit learning or produce effects similar to studying with fixed spacing."
They name the software vendors directly, noting that this absence of evidence "has not stopped some software developers from assuming that expanding study intervals work better than fixed intervals," and citing SuperMemo's universal formula as the example.
So: I ship an expanding schedule, and the strongest meta-analysis in this literature says expanding is not demonstrably better than fixed. I keep it for reasons that are practical rather than empirical - an expanding schedule spends less of your attention on a mock you have already squeezed dry, which matters when you are running twenty of them - but you should know that the schedule's shape is a design choice and only its existence is evidence-backed. Anyone selling you a specific interval sequence as science is going beyond what Cepeda et al. found.
What my own 20 mocks show about gaps
Twenty mocks between 10 May and 8 August 2026 gives 19 gaps. Here is the distribution.
| Gap between consecutive mocks | Count | What it produced |
|---|---|---|
| 1-2 days | 7 | Includes the 96 on 29 July and the 40 on 30 July |
| 5-6 days | 4 | Enough room for one analysis pass |
| 7 days | 8 | The weekly coaching cadence, and the useful one |
Median gap: 6 days. Mean: 4.7 days, pulled down by the clusters. The seven gaps of two days or less are the ones I would remove if I ran the quarter again - not because sitting a mock is harmful, but because a mock taken 48 hours after another one cannot be revised in between, so it measures your state rather than teaching you anything. My clearest example is public: I scored 96 on 29 July and 40 the next day, with 2 marks in DILR. Nothing about my knowledge changed overnight. The full log is in my 20-mock post.
What I would actually do
Grounded in the above, not in vibes:
- Two revision passes minimum per mock, days apart. The first pass is still emotionally attached to the score. Spacing is what lets the second pass see something new.
- Put the first pass 2-3 days out, not the same evening. The evening pass is massed practice with extra steps. Cepeda's data cannot resolve 1 day from 3 days, so this is a judgement call, but it is on the right side of the one thing the paper is confident about.
- Space the later passes further apart than the earlier ones, and do not believe anyone who tells you the exact numbers matter. Fixed would probably work as well. Expanding is easier to sustain.
- Do not sit two mocks inside 48 hours unless you are deliberately simulating back-to-back pressure. Seven of my nineteen gaps broke this rule and none of them produced a usable signal.
- Count revision passes, not mocks. Twenty mocks revised once is worse than twelve revised three times. See how many mocks are actually enough.
Where this evidence does not reach
I would rather lose your click than overstate this, so here is the full list of reasons the numbers above should not be treated as a CAT prescription.
- The materials are word lists, not DILR sets. The title says it: "distributed practice in verbal recall tasks". Paired associates and list recall. Learning that a spaced word list is remembered better does not establish that a spaced re-analysis of a logical-reasoning set improves your set selection under a 40-minute clock.
- 85% of the data are young adults in a laboratory. Cepeda et al. state this explicitly. CAT is sat by roughly 2.5 lakh Indian graduates in a three-hour proctored window with sectional time locks and negative marking. The mechanism plausibly transfers; the 10.6-point gap absolutely does not transfer as a number.
- The retention intervals that match your situation are nearly unstudied. One study, 43 participants, for anything beyond 31 days. The authors say new studies at educationally relevant intervals "are sorely needed", and that was twenty years ago.
- Expanding versus fixed is unresolved, as above, and my product picks a side anyway.
- My own log is one candidate. Twenty mocks, one person, thirteen weeks, no actual CAT score to check any of it against - I have not sat CAT yet. It is enough to show what a clustered mock schedule looks like. It is not enough to establish that a 6-day gap beats a 3-day one.
- Nobody has run the study you want. There is no randomised trial of mock-revision spacing in Indian competitive exams. If one exists, I have not found it, and I looked.
Sources
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3). Read in full at eScholarship. All figures above are from its Tables 1 and 8 and its Discussion.
- Mock log, 20 attempts, 10 May to 8 August 2026 - first-party, published in full in the 20-mock post.
The revision rounds described here are the ones Karma Yogi tracks against each logged mock, so the gaps are recorded rather than remembered. Whether the exact intervals are optimal is, on the evidence above, an open question.
End of essay
- Anish Guruvelli