Karma YogiKarma Yogi
Start
The Journal
CAT Preparation

10 September 2026
12 min read

A
Anish Guruvelli
Karma Yogi

Deliberate Practice for CAT: What 88 Studies Actually Say

The "quality beats quantity" claim has a real evidence base and a much smaller effect than the popular version admits: 12% of performance variance overall, 4% in education. Here is what survives contact with the meta-analysis, and what my own DILR log says about practising a section for three months and moving two marks.

Practice quality beats practice quantity - but by less than the popular version of that claim promises, and least of all in the domain you are actually in. The strongest test is Macnamara, Hambrick and Oswald's 2014 meta-analysis in Psychological Science, which pooled 88 studies, 157 effect sizes and 11,135 participants and found accumulated deliberate practice explained 12% of the variance in performance overall, and 4% in education. Against that, my own log of 20 CAT mocks between 10 May and 8 August 2026: my DILR average moved from 18.4 across the first five mocks to 20.6 across the last five. Two marks, in thirteen weeks, while VARC gained 8.2 and Quant gained 8.8.

That is the honest headline, and both halves of it matter. Practice type is worth optimising. It is not the whole answer, nobody serious claims a number as high as the internet does, and the evidence is weakest exactly where a CAT aspirant sits.

What Ericsson actually said, and on what evidence

The idea comes from Ericsson, Krampe and Tesch-Römer (1993) in Psychological Review. Their definition is narrower than the way the phrase gets used, and worth reading in their own words:

"In contrast to play, deliberate practice is a highly structured activity, the explicit goal of which is to improve performance. Specific tasks are invented to overcome weaknesses, and performance is carefully monitored to provide cues for ways to improve it further. We claim that deliberate practice requires effort and is not inherently enjoyable."

Four conditions are doing the work there: a specific weakness, a task invented to attack it, performance monitored, and effort that is not enjoyable. Solving 40 more Quant questions from the same book is not deliberate practice under that definition. It is repetition, which is a different thing with a different effect size.

Now the part that rarely gets quoted. That framework rests on two studies. Study 1 was 40 violinists at one Berlin music academy - three matched groups of ten young players plus ten middle-aged orchestral professionals. Study 2 was 24 pianists, 12 expert and 12 amateur. Sixty-four people, one city, one instrument family. The paper is a landmark and deserves to be; it is also a very small foundation for a claim that has since been applied to every skill anybody wants to sell a course in.

The meta-analysis that reopened it

Macnamara et al. (2014) collected every study they could find that reported a correlation between accumulated deliberate practice and performance - 88 studies, 157 effect sizes, 11,135 participants. The mean correlation was r = .35, 95% CI [.30, .39], which is 12% of variance explained and 88% unexplained. Then they split it by domain, and the split is the finding.

Domain Variance explained Correlation Significance
Games (chess, Scrabble)26%.51p < .001
Music21%.46p < .001
Sports18%.42p < .001
Education4%.21p < .001
Professions< 1%.05p = .62 (not significant)
All domains pooled12%.35p < .001

Domain was a significant moderator, Q(4) = 49.09, p < .001. Education is the fourth row. CAT prep is not chess and it is not violin - it is closest to the row where the effect is 4%, and it is the row the authors themselves flag as poorly defined:

"Why were the effect sizes for education and professions so much smaller? One possibility is that deliberate practice is less well defined in these domains. It could also be that in some of the studies, participants differed in amount of prestudy expertise... and thus in the amount of deliberate practice they needed to achieve a given level of performance."

That second clause is the CAT-relevant one. An engineer starting CAT prep already carries a decade of Quant fluency; a humanities graduate does not. The two need different amounts of practice to reach the same score, which mechanically weakens any correlation between hours practised and score achieved across a mixed cohort. It does not mean practice is not working for you. It means the cross-sectional correlation is the wrong instrument for answering that question about you.

They also tested a second moderator that reads almost like a description of the CAT: predictability of the task environment. Effects were largest (24%) for highly predictable activities such as running, intermediate (12%) for moderately predictable ones, and smallest (4%) for low-predictability activities. A DILR set you have never seen before, under a 40-minute lock, is a low-predictability environment by construction.

The finding that should change what you do tomorrow

The methodological moderator is the most actionable result in the paper and gets almost no coverage. How the study measured practice changed the answer by a factor of four.

How practice was measured Variance explained Correlation
Retrospective interview20%.45
Retrospective questionnaire12%.34
Log (diary or computer, recorded as it happened)5%.22

Q(2) = 16.19, p < .001. The more accurately practice was recorded, the smaller the apparent effect. Macnamara et al. read that the obvious way: retrospective recall inflates the relationship, and "for studies using the log method, which presumably yields more valid estimates than retrospective methods do, deliberate practice accounted for only 5% of the variance."

Here is the detail that closes the loop, and it comes from Ericsson's own 1993 paper rather than his critics'. He had his violinists estimate a typical week's practice, then keep a 7-day diary. The estimates came in 5.2 hours per week higher than the diaries, F(1, 27) = 15.39, p < .001 - and the same 5.2-hour gap replicated in the pianist study. His explanation:

"Debriefing interviews suggested that the estimates for a typical week reflected a level of daily practice to which the violinists aspired rather than the level they actually attained."

Elite conservatory musicians, asked about the single activity their entire life is organised around, overstated it by five hours a week. If you are estimating your own CAT hours at the end of a month, you are not doing better than they did. The practical takeaway is not "practise more deliberately." It is "stop estimating." Everything else in this post is downstream of having a real record rather than an aspirational one.

What deliberate practice looks like when you translate it to CAT

Ericsson's four conditions are testable against a study session. Most CAT prep fails at least two of them. This table is my own translation, not a finding - it is what the definition implies when you hold it against the things aspirants actually do.

Activity Specific weakness? Task built for it? Performance monitored? Verdict
Taking a full mockNoNoYesMeasurement, not practice
Re-reading a solved solutionSometimesNoNoWeakest common habit
40 mixed Quant questions from a bookNoNoPartlyVolume, not deliberate
Ten arrangement sets, timed, one grid techniqueYesYesYesMeets the definition
Re-attempting only the RC questions you got wrong, blindYesYesYesMeets the definition
Set-selection drill: pick 3 of 4 sets in 4 minutes, never solveYesYesYesMeets it, and nobody does it

The last row is the one I wish I had run for three months. Set selection is the decision that determines a DILR score, it is separable from solving, and it can be drilled in four-minute blocks. Instead I did what the third row describes.

What my own log says, and it is not flattering

Twenty mocks, 10 May to 8 August 2026. Comparing the first five attempts to the last five:

Section First 5 mocks (avg) Last 5 mocks (avg) Change
VARC16.624.8+8.2
QUANT26.835.6+8.8
DILR18.420.6+2.2
Overall (of 198)61.881.0+19.2

I logged real hours against DILR across that quarter and moved 2.2 marks. I cannot run the counterfactual - there is no version of me who spent the same hours on set-selection drills instead - so this is one candidate's experience, not evidence. What it does illustrate is the failure mode the definition predicts: the two sections where my practice had a specific target moved; the one where I was mostly re-solving whole sets did not. The full log, crashes included, is in my 20-mock post.

What I would actually do

  • Log as you go, not at the end of the week. This is the one recommendation the evidence supports directly, from both directions: the log method is what shrinks the inflated effect, and Ericsson's own subjects were 5.2 hours a week optimistic about themselves.
  • Name the weakness before the session, not after. "Two hours of DILR" fails the definition. "Ten arrangement sets, 4 minutes each, no solving past the grid" passes it.
  • Separate measurement from practice. A mock measures. It does not train. Treat the number of mocks and the number of drills as two different counters - see how many mocks are actually enough.
  • Drill the decision, not only the solution. Set selection, question skipping and abandonment points are all separable skills that almost nobody practises in isolation, and they are where a low-predictability section is won.
  • Do not conclude from a flat section that you lack talent. 88% of the variance in that meta-analysis was unexplained by practice - by everything else, of which innate ability is one candidate among many, including starting point, method, sleep and what you were actually doing in those hours.

Where this is weak, and where I am arguing against my own framing

  • "12% of variance" is not "practice barely matters". That misreading is as wrong as the 10,000-hour one. Of 157 correlations, nearly all were positive and only two were significantly negative. Macnamara et al.'s own conclusion is that deliberate practice is "unquestionably important as a predictor... but not as important as Ericsson and his colleagues have argued." Both clauses are load-bearing.
  • None of these studies are about CAT. The education bucket is coursework and academic performance, not a three-hour proctored aptitude exam with sectional locks and negative marking. I am reasoning from an adjacent domain and saying so.
  • Correlational, not causal. Every study in the meta-analysis measures accumulated practice against attained performance. People who are already good practise more, and get better feedback while doing it. Nothing here randomly assigns practice hours.
  • Heterogeneity was high. I² = 84.90, meaning most of the spread between studies is real difference rather than sampling noise. A pooled 12% is a summary of a very wide distribution, not a constant.
  • Uncorrected for measurement unreliability. The authors flag this and give the corrected figures: assuming .80 reliability on both sides, the overall correlation rises to .43 - 19% of reliable variance, and 7% for education. Higher than 12% and 4%. Still not most of it.
  • The definition is easy to abuse in one's own favour. "That was deliberate practice" is unfalsifiable after the fact, which is precisely why the four conditions have to be written down before the session rather than argued for afterwards.
  • My DILR result proves nothing on its own. One person, twenty mocks, thirteen weeks, no CAT score to check any of it against - I have not sat the exam yet.

Sources

  • Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance in music, games, sports, education, and professions: A meta-analysis. Psychological Science, 25(8), 1608-1618. doi:10.1177/0956797614535810. All domain, predictability and measurement-method figures above are from its Results and Discussion.
  • Ericsson, K. A., Krampe, R. Th., & Tesch-Römer, C. (1993). The role of deliberate practice in the acquisition of expert performance. Psychological Review, 100(3), 363-406. doi:10.1037/0033-295X.100.3.363. The definition, the sample sizes and the 5.2-hour estimate-versus-diary gap are from its Studies 1 and 2.
  • Mock log, 20 attempts, 10 May to 8 August 2026 - first-party, published in full in the 20-mock post.

The logging half of this is the part Karma Yogi exists to make automatic: sessions recorded as they happen, tagged to a subject and a topic, so the number you look back on in November is a diary rather than an aspiration. Whether any given session was deliberate is still a judgement only you can make before you start it.

End of essay

- Anish Guruvelli

Common questions

Does the 10,000-hour rule apply to CAT preparation?
No, and it does not really apply anywhere in the form it is usually quoted. The figure comes from Ericsson et al.'s 1993 study of 40 violinists and 24 pianists at one Berlin academy, describing an average for elite performers rather than a threshold. For education specifically, a 2014 meta-analysis of 88 studies found accumulated practice explained only 4% of performance variance.
What actually counts as deliberate practice?
Ericsson's definition has four parts: you are attacking a specific weakness, using a task designed for that weakness, with performance monitored for corrective feedback, and it is effortful rather than enjoyable. Solving another 40 mixed questions from a book fails at least two of them. Ten timed arrangement sets targeting one grid technique passes all four.
Is quality of study really better than quantity?
Better, but not by the margin the phrase implies, and the honest version is that both matter and neither explains most of the outcome. Across 88 studies and 11,135 participants, deliberate practice explained 12% of performance variance overall, leaving 88% to everything else - method, starting point, ability, sleep, and what you were actually doing during those hours.
Why did my mock scores stop improving even though I keep practising?
Usually because the practice stopped being targeted, not because you stopped being capable. My own DILR average moved 2.2 marks across 20 mocks and thirteen weeks while VARC gained 8.2 and Quant gained 8.8, and the difference was that the other two sections had a specific weakness named before each session. A plateau is a signal to change the task, not to add hours.
Should I take more mocks or do more targeted drills?
Drills, in most cases, because a mock is a measurement instrument rather than a training one. Under the deliberate-practice definition a full mock fails three of the four conditions: no specific weakness, no task built for it, and nothing isolated to correct. Take enough mocks to measure and diagnose, then spend the bulk of your hours on what the diagnosis named.
Does tracking my study hours actually improve anything?
It improves the accuracy of every decision you make from the number, which is more than it sounds. In Ericsson's own 1993 data, elite violinists overestimated their weekly practice by 5.2 hours compared with their own diaries, and he attributed the gap to reporting the level they aspired to. Estimates drift upward; a log does not.
Why is deliberate practice less effective in education than in chess or music?
Two reasons the 2014 meta-analysis gives. Deliberate practice is less clearly defined in education, so studies are measuring a fuzzier thing. And students arrive with very different amounts of prior knowledge, so they need different amounts of practice to reach the same score - which mechanically weakens any correlation between hours and outcome across a mixed group.
Is DILR harder to improve than VARC or Quant?
It behaves differently rather than being intrinsically harder, and the meta-analysis suggests why. Effects of practice were smallest, around 4% of variance, for low-predictability activities, and an unseen DILR set under a sectional lock is about as unpredictable as exam tasks get. Practising the decision - which sets to pick, when to abandon one - is more tractable than practising the solving.
How many hours a day should I study for CAT?
No study in this literature can answer that, and anyone quoting a specific number is going beyond the evidence. What the research does support is that how the hours are structured matters at least as much as their count, and that your own estimate of how many you did is likely to be optimistic unless you recorded them as they happened.
Is talent real, or is it all practice?
The 2014 meta-analysis leaves 88% of performance variance unexplained by practice, and that residual contains many things - starting age, prior knowledge, general cognitive ability, method, motivation, health - not just innate talent. The useful reading is not "talent decides it" but "practice is one of several inputs, and it is the one you control".
Does re-reading solutions to questions I got wrong help?
It is the weakest of the common habits, because it fails the monitoring condition: you are reading a correct answer rather than generating one, so nothing measures whether you could now do it. Re-attempting the same questions blind, days later, meets the definition and takes the same amount of time.
How do I tell whether a study session was deliberate practice?
Decide before it starts, not after. Write down the specific weakness, the task built to attack it, and how you will know whether it worked. If you can only tell yourself afterwards that a session was deliberate, the label is unfalsifiable and probably wrong.