Close the solution and reconstruct it. Do not read it again. In Karpicke and Blunt's 2011 experiment in Science, 101 of 120 students performed better on a delayed test after practising retrieval than after an elaborative study method - and only 30 of those 120 predicted they would. My own reason for caring is narrower: across five consecutive CAT mocks my VARC scores went 3, 36, 5, 39, 9, and reading passages again was most of what I was doing about it.
The gap between the 101 and the 30 is the entire problem. Rereading survives not because anyone has evidence for it but because it produces a strong feeling of fluency, and the technique that actually works feels worse while you are doing it.
The study, in enough detail to check
Jeffrey Karpicke and Janell Blunt, Purdue University, published in Science in February 2011. Two experiments.
Experiment 1: 80 undergraduates. Everyone studied a science text, then split into four conditions:
- Study once - a single reading.
- Repeated study - four consecutive study periods. This is rereading, formalised.
- Elaborative concept mapping - study the text, then build a concept map while looking at it. An active, effortful, widely taught study method.
- Retrieval practice - study the text, then write down as much as you could recall with the text closed, restudy, recall again.
Total learning time was matched exactly between the concept-mapping and retrieval groups. One week later everyone sat a short-answer test containing both verbatim questions and inference questions that required connecting concepts across the text.
| Experiment | Comparison | Retrieval | Elaborative study | Effect size |
|---|---|---|---|---|
| 1 (n = 80) | Short-answer test, 1 week later | 0.67 | 0.45 | d = 1.50 |
| 2 (n = 120) | Short-answer test, 1 week later | - | - | d = 1.07 |
| 2 (n = 120) | Final test was building a concept map | - | - | d = 1.01 |
The Experiment 1 gap - 0.67 against 0.45 - is described by the authors as "about a 50% improvement in long-term retention scores", at d = 1.50, F(1,38) = 21.63.
Two details in that experiment matter more than the headline. First, concept mapping was not significantly better than simply spending more time reading. An effortful, active, respectable study method performed no better than rereading. Second, in Experiment 2 the final test for half the students was building a concept map from memory - the format that should have favoured the mapping group - and retrieval practice still won, d = 1.01.
The number that explains why nobody does this
Experiment 2 used a within-subject design: each of the 120 students concept-mapped one text and practised retrieval on another, so you can count individuals rather than group means. The authors also asked each student to predict how much they would remember.
| Out of 120 students | Retrieval better | No difference | Elaborative study better |
|---|---|---|---|
| What actually happened | 101 | 6 | 13 |
| What students predicted | 30 | 31 | 59 |
84% of students learned more from retrieval practice. 75% believed the opposite method would be as good or better. The authors are explicit that in Experiment 1 the ranking students predicted was almost exactly inverted: they expected repeated studying to produce the best retention and practising retrieval to produce the worst.
This is not a one-off. Samani and Pan's 2021 study of 350 undergraduates in a UCLA physics course found the same shape from a different angle - students rated the more effective practice format as more difficult and believed they had learned less from it, while scoring 50% to 125% higher on the delayed tests. Across two very different experiments, the technique that works is the one that feels worse in the moment.
Why this indicts the standard mock-analysis workflow
Here is what almost everyone does after a CAT mock. Open the solution PDF. Find the questions you got wrong. Read the explanation. Think "ah, of course". Move on.
That is elaborative studying with the material in front of you, which is precisely the condition that lost in both experiments. It is the concept-mapping group: active, effortful, feels productive, and in Karpicke and Blunt's data no better than rereading.
The retrieval-practice version of the same 40 minutes looks like this:
- Re-attempt before you read anything. Take the questions you got wrong, close everything, and solve them cold. The retrieval attempt is the learning event, and it works even when you fail it.
- Write the method from memory before checking. Not the answer - the approach. If you cannot state why the correct option is correct without looking, you have not learned it, whatever the feeling of recognition says.
- Only then open the solution, and only to correct the specific gap you just proved you had.
- Come back days later and do it again. Retrieval and spacing compound; the strongest version of this is a second cold attempt a week on. See what the spacing meta-analysis actually says for how far that evidence goes.
The uncomfortable part is that the second workflow will feel worse. You will get things wrong out loud, in writing, twice. That feeling is not a signal that it is not working - in this literature it is close to a signal that it is.
What my own log looks like under this lens
VARC was my most volatile section across 20 mocks between May and August 2026. The sequence across five consecutive papers:
| Date | Mock | VARC score | Percentile |
|---|---|---|---|
| 10 May | PreSimCAT 03 | 3 | 12 |
| 20 Jun | SIMCAT 4 | 36 | 92 |
| 28 Jun | SIMCAT 104 | 5 | 15 |
| 11 Jul | SIMCAT 105 | 39 | 78 |
| 27 Jul | SIMCAT 7 | 9 | 37 |
There is no learning curve in that. It is close to noise, and the 5 on 28 June came down to a single distortion I would catch on a calm day: the passage discussed only Descartes and the option said "Descartes and other philosophers". I took it.
What I was doing about VARC at the time was reading passages and reading solutions. What I was not doing was reconstructing, from memory and without the passage, why a wrong option was wrong. The section that did move steadily over the same quarter - QUANT, from a 26.8 average across my first five mocks to 35.6 across my last five - is the one where I was solving problems cold, because that is the only way quant practice works. The section I effectively reread was the section that stayed random.
I cannot prove causation from one candidate's log. I can say the pattern is exactly what the experiments above would predict, and it is the reason I changed how I analyse a paper.
Where this evidence does not reach
- d = 1.50 will not survive contact with your kitchen table. That is an enormous effect from a controlled laboratory session with matched learning time, one week of delay, and one text. Classroom effects are routinely smaller than laboratory ones, and sometimes vanish entirely - the same UCLA physics study cited above found a large benefit on its surprise tests and no significant benefit on the students' actual midterm exams. Expect the direction to hold and the magnitude to shrink.
- The task is not your task. Karpicke and Blunt used science prose and free recall of its content. CAT asks you to select and solve unfamiliar problems against a clock with negative marking. Recalling a text and choosing a DILR set are not the same cognitive act, and no study here tested the second one.
- The population is 200 American undergraduates. Not 2.5 lakh Indian graduates, most of them working, revising over months rather than one week.
- The obvious follow-up question is one I cannot answer from a source I have read. Whether test-enhanced learning transfers to different questions from the ones you practised is the subject of a 2018 meta-analysis in Psychological Bulletin. It is paywalled, I have not read it, and so nothing in this post rests on it. If you see a study-tips post confidently quoting a transfer effect size, ask where they read it.
- 13 of 120 students did worse with retrieval practice. That is about 11%. This is a strong average effect, not a law, and if you have honestly tried the closed-book loop for a month and your scores are flat, the data allows for you being in the 13.
- My VARC data is one person, five mocks, and no controlled comparison. It is an illustration of the argument, not evidence for it.
Sources
- Karpicke, J. D., & Blunt, J. R. (2011). Retrieval practice produces more learning than elaborative studying with concept mapping. Science, 331(6018), 772-775. Read in full from the authors' own publications archive. All figures above are from its Figures 1-2 and Table 1.
- Samani, J., & Pan, S. C. (2021). Interleaved practice enhances memory and problem-solving ability in undergraduate physics. npj Science of Learning, 6:32. Open access. Cited here only for its student-perception and midterm findings.
- Mock log, 20 attempts, 10 May to 8 August 2026 - first-party, published in full in the 20-mock post.
Karma Yogi logs each analysis pass against the mock it belongs to, which is the only way I found to tell a second cold attempt apart from a second read of the same solution. They feel identical at the time and they are not the same thing.
End of essay
- Anish Guruvelli