Practice quality beats practice quantity - but by less than the popular version of that claim promises, and least of all in the domain you are actually in. The strongest test is Macnamara, Hambrick and Oswald's 2014 meta-analysis in Psychological Science, which pooled 88 studies, 157 effect sizes and 11,135 participants and found accumulated deliberate practice explained 12% of the variance in performance overall, and 4% in education. Against that, my own log of 20 CAT mocks between 10 May and 8 August 2026: my DILR average moved from 18.4 across the first five mocks to 20.6 across the last five. Two marks, in thirteen weeks, while VARC gained 8.2 and Quant gained 8.8.
That is the honest headline, and both halves of it matter. Practice type is worth optimising. It is not the whole answer, nobody serious claims a number as high as the internet does, and the evidence is weakest exactly where a CAT aspirant sits.
What Ericsson actually said, and on what evidence
The idea comes from Ericsson, Krampe and Tesch-Römer (1993) in Psychological Review. Their definition is narrower than the way the phrase gets used, and worth reading in their own words:
"In contrast to play, deliberate practice is a highly structured activity, the explicit goal of which is to improve performance. Specific tasks are invented to overcome weaknesses, and performance is carefully monitored to provide cues for ways to improve it further. We claim that deliberate practice requires effort and is not inherently enjoyable."
Four conditions are doing the work there: a specific weakness, a task invented to attack it, performance monitored, and effort that is not enjoyable. Solving 40 more Quant questions from the same book is not deliberate practice under that definition. It is repetition, which is a different thing with a different effect size.
Now the part that rarely gets quoted. That framework rests on two studies. Study 1 was 40 violinists at one Berlin music academy - three matched groups of ten young players plus ten middle-aged orchestral professionals. Study 2 was 24 pianists, 12 expert and 12 amateur. Sixty-four people, one city, one instrument family. The paper is a landmark and deserves to be; it is also a very small foundation for a claim that has since been applied to every skill anybody wants to sell a course in.
The meta-analysis that reopened it
Macnamara et al. (2014) collected every study they could find that reported a correlation between accumulated deliberate practice and performance - 88 studies, 157 effect sizes, 11,135 participants. The mean correlation was r = .35, 95% CI [.30, .39], which is 12% of variance explained and 88% unexplained. Then they split it by domain, and the split is the finding.
| Domain | Variance explained | Correlation | Significance |
|---|---|---|---|
| Games (chess, Scrabble) | 26% | .51 | p < .001 |
| Music | 21% | .46 | p < .001 |
| Sports | 18% | .42 | p < .001 |
| Education | 4% | .21 | p < .001 |
| Professions | < 1% | .05 | p = .62 (not significant) |
| All domains pooled | 12% | .35 | p < .001 |
Domain was a significant moderator, Q(4) = 49.09, p < .001. Education is the fourth row. CAT prep is not chess and it is not violin - it is closest to the row where the effect is 4%, and it is the row the authors themselves flag as poorly defined:
"Why were the effect sizes for education and professions so much smaller? One possibility is that deliberate practice is less well defined in these domains. It could also be that in some of the studies, participants differed in amount of prestudy expertise... and thus in the amount of deliberate practice they needed to achieve a given level of performance."
That second clause is the CAT-relevant one. An engineer starting CAT prep already carries a decade of Quant fluency; a humanities graduate does not. The two need different amounts of practice to reach the same score, which mechanically weakens any correlation between hours practised and score achieved across a mixed cohort. It does not mean practice is not working for you. It means the cross-sectional correlation is the wrong instrument for answering that question about you.
They also tested a second moderator that reads almost like a description of the CAT: predictability of the task environment. Effects were largest (24%) for highly predictable activities such as running, intermediate (12%) for moderately predictable ones, and smallest (4%) for low-predictability activities. A DILR set you have never seen before, under a 40-minute lock, is a low-predictability environment by construction.
The finding that should change what you do tomorrow
The methodological moderator is the most actionable result in the paper and gets almost no coverage. How the study measured practice changed the answer by a factor of four.
| How practice was measured | Variance explained | Correlation |
|---|---|---|
| Retrospective interview | 20% | .45 |
| Retrospective questionnaire | 12% | .34 |
| Log (diary or computer, recorded as it happened) | 5% | .22 |
Q(2) = 16.19, p < .001. The more accurately practice was recorded, the smaller the apparent effect. Macnamara et al. read that the obvious way: retrospective recall inflates the relationship, and "for studies using the log method, which presumably yields more valid estimates than retrospective methods do, deliberate practice accounted for only 5% of the variance."
Here is the detail that closes the loop, and it comes from Ericsson's own 1993 paper rather than his critics'. He had his violinists estimate a typical week's practice, then keep a 7-day diary. The estimates came in 5.2 hours per week higher than the diaries, F(1, 27) = 15.39, p < .001 - and the same 5.2-hour gap replicated in the pianist study. His explanation:
"Debriefing interviews suggested that the estimates for a typical week reflected a level of daily practice to which the violinists aspired rather than the level they actually attained."
Elite conservatory musicians, asked about the single activity their entire life is organised around, overstated it by five hours a week. If you are estimating your own CAT hours at the end of a month, you are not doing better than they did. The practical takeaway is not "practise more deliberately." It is "stop estimating." Everything else in this post is downstream of having a real record rather than an aspirational one.
What deliberate practice looks like when you translate it to CAT
Ericsson's four conditions are testable against a study session. Most CAT prep fails at least two of them. This table is my own translation, not a finding - it is what the definition implies when you hold it against the things aspirants actually do.
| Activity | Specific weakness? | Task built for it? | Performance monitored? | Verdict |
|---|---|---|---|---|
| Taking a full mock | No | No | Yes | Measurement, not practice |
| Re-reading a solved solution | Sometimes | No | No | Weakest common habit |
| 40 mixed Quant questions from a book | No | No | Partly | Volume, not deliberate |
| Ten arrangement sets, timed, one grid technique | Yes | Yes | Yes | Meets the definition |
| Re-attempting only the RC questions you got wrong, blind | Yes | Yes | Yes | Meets the definition |
| Set-selection drill: pick 3 of 4 sets in 4 minutes, never solve | Yes | Yes | Yes | Meets it, and nobody does it |
The last row is the one I wish I had run for three months. Set selection is the decision that determines a DILR score, it is separable from solving, and it can be drilled in four-minute blocks. Instead I did what the third row describes.
What my own log says, and it is not flattering
Twenty mocks, 10 May to 8 August 2026. Comparing the first five attempts to the last five:
| Section | First 5 mocks (avg) | Last 5 mocks (avg) | Change |
|---|---|---|---|
| VARC | 16.6 | 24.8 | +8.2 |
| QUANT | 26.8 | 35.6 | +8.8 |
| DILR | 18.4 | 20.6 | +2.2 |
| Overall (of 198) | 61.8 | 81.0 | +19.2 |
I logged real hours against DILR across that quarter and moved 2.2 marks. I cannot run the counterfactual - there is no version of me who spent the same hours on set-selection drills instead - so this is one candidate's experience, not evidence. What it does illustrate is the failure mode the definition predicts: the two sections where my practice had a specific target moved; the one where I was mostly re-solving whole sets did not. The full log, crashes included, is in my 20-mock post.
What I would actually do
- Log as you go, not at the end of the week. This is the one recommendation the evidence supports directly, from both directions: the log method is what shrinks the inflated effect, and Ericsson's own subjects were 5.2 hours a week optimistic about themselves.
- Name the weakness before the session, not after. "Two hours of DILR" fails the definition. "Ten arrangement sets, 4 minutes each, no solving past the grid" passes it.
- Separate measurement from practice. A mock measures. It does not train. Treat the number of mocks and the number of drills as two different counters - see how many mocks are actually enough.
- Drill the decision, not only the solution. Set selection, question skipping and abandonment points are all separable skills that almost nobody practises in isolation, and they are where a low-predictability section is won.
- Do not conclude from a flat section that you lack talent. 88% of the variance in that meta-analysis was unexplained by practice - by everything else, of which innate ability is one candidate among many, including starting point, method, sleep and what you were actually doing in those hours.
Where this is weak, and where I am arguing against my own framing
- "12% of variance" is not "practice barely matters". That misreading is as wrong as the 10,000-hour one. Of 157 correlations, nearly all were positive and only two were significantly negative. Macnamara et al.'s own conclusion is that deliberate practice is "unquestionably important as a predictor... but not as important as Ericsson and his colleagues have argued." Both clauses are load-bearing.
- None of these studies are about CAT. The education bucket is coursework and academic performance, not a three-hour proctored aptitude exam with sectional locks and negative marking. I am reasoning from an adjacent domain and saying so.
- Correlational, not causal. Every study in the meta-analysis measures accumulated practice against attained performance. People who are already good practise more, and get better feedback while doing it. Nothing here randomly assigns practice hours.
- Heterogeneity was high. I² = 84.90, meaning most of the spread between studies is real difference rather than sampling noise. A pooled 12% is a summary of a very wide distribution, not a constant.
- Uncorrected for measurement unreliability. The authors flag this and give the corrected figures: assuming .80 reliability on both sides, the overall correlation rises to .43 - 19% of reliable variance, and 7% for education. Higher than 12% and 4%. Still not most of it.
- The definition is easy to abuse in one's own favour. "That was deliberate practice" is unfalsifiable after the fact, which is precisely why the four conditions have to be written down before the session rather than argued for afterwards.
- My DILR result proves nothing on its own. One person, twenty mocks, thirteen weeks, no CAT score to check any of it against - I have not sat the exam yet.
Sources
- Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance in music, games, sports, education, and professions: A meta-analysis. Psychological Science, 25(8), 1608-1618. doi:10.1177/0956797614535810. All domain, predictability and measurement-method figures above are from its Results and Discussion.
- Ericsson, K. A., Krampe, R. Th., & Tesch-Römer, C. (1993). The role of deliberate practice in the acquisition of expert performance. Psychological Review, 100(3), 363-406. doi:10.1037/0033-295X.100.3.363. The definition, the sample sizes and the 5.2-hour estimate-versus-diary gap are from its Studies 1 and 2.
- Mock log, 20 attempts, 10 May to 8 August 2026 - first-party, published in full in the 20-mock post.
The logging half of this is the part Karma Yogi exists to make automatic: sessions recorded as they happen, tagged to a subject and a topic, so the number you look back on in November is a diary rather than an aspiration. Whether any given session was deliberate is still a judgement only you can make before you start it.
End of essay
- Anish Guruvelli