A plateau in your overall mock score is almost never a plateau in your ability - it is usually two real trends cancelling, and the fix is a change in what you practise rather than more of it. In my own log of 20 CAT mocks, the mean overall score across three consecutive blocks of five went 61.8, then 59.8, then 62.0. Ten weeks, no movement. Underneath that flat line, my VARC mean rose from 16.6 to 24.4 while my QUANT mean fell from 26.8 to 23.4. Two genuine trends, pointed in opposite directions, summing to nothing. The research says the same thing at a much larger scale: an analysis of 25,280 individual learning curves found that the smooth improvement curve everybody quotes is largely an artefact of averaging, and that real individual curves contain plateaus, dips and sudden jumps.
So "my mock scores are not improving" is usually the wrong diagnosis. The right question is which of your three sections is moving, in which direction, and what you would have to change to stop them cancelling.
My actual plateau, in blocks of five
Twenty full mocks between 10 May and 8 August 2026, grouped in consecutive blocks of five, scores out of 198. The last three attempts in block four are real CAT 2020 papers rather than coaching mocks - a confound I deal with properly further down rather than burying.
| Block | Overall (mean) | Percentile (mean) | VARC | DILR | QUANT |
|---|---|---|---|---|---|
| Mocks 1-5 | 61.8 | 73.4 | 16.6 | 18.4 | 26.8 |
| Mocks 6-10 | 59.8 | 80.2 | 17.0 | 19.2 | 23.6 |
| Mocks 11-15 | 62.0 | 80.5 | 24.4 | 14.2 | 23.4 |
| Mocks 16-20 | 81.0 | 94.3 | 24.8 | 20.6 | 35.6 |
Read the middle two rows. Ten weeks and ten mocks in which the number I was actually looking at - the overall score - moved by 2.2 marks in the wrong direction and then 2.2 back. If you had asked me in mid-July whether I was improving, I would have said no, and I would have been wrong in a specific and useful way.
The plateau was three sections cancelling
Between block one and block three, with the overall flat:
- VARC: 16.6 to 24.4. Up 7.8 marks. This was the daily reading habit and the RC question-type work finally landing.
- QUANT: 26.8 to 23.4. Down 3.4 marks. My strongest section going backwards while I paid it no attention, precisely because it was my strongest.
- DILR: 18.4 to 14.2. Down 4.2 marks, and this was the real problem hiding under the flat line.
Add those up and you get roughly zero, which is exactly what the overall column shows. An aggregate that is not moving is not evidence that nothing is moving. It is a single number doing what single numbers do - averaging away the information you need.
This is the same failure mode, one level down, that the learning-curve literature describes at population scale, and I find that symmetry genuinely useful: whether you average across three sections or across twenty-two thousand people, averaging manufactures smoothness that is not in the underlying data.
What the learning-curve research actually says about plateaus
Donner and Hardy, Psychonomic Bulletin & Review 22:1308-1319, 2015. They analysed 25,280 learning curves from 22,460 people on the Lumosity cognitive training platform, each curve carrying 500 measurements of performance on the same task. That is an unusually good look at what individual learning actually looks like, because almost every classic study of practice averaged across participants first and drew a curve second.
Three findings that change how you should read your own mock log:
- Individual curves are not smooth. A piecewise power law - a curve made of discrete segments with breakpoints - explained 90.74% of variance against 85.97% for a single smooth power law, a difference so large it is hard to overstate statistically (chi-squared (94205) = 534,030). Two- and three-segment solutions were the most common.
- Real curves contain "both plateaus and bursts of rapid improvement". The smooth curve, the authors write, "may be an artifact of averaging across individual curves". The tidy diminishing-returns picture you have seen is a population statistic, not a description of anybody.
- The segments are strategy changes. They describe the pattern as "a discrete sequence of strategy shifts, in which each strategy is better in the long term than the ones preceding it".
That third point is the practically important one and it is the whole argument of this post. If a plateau is the tail of one strategy rather than a limit on your ability, then more repetitions of that strategy is the one intervention guaranteed not to work. You do not out-practise a plateau. You change what is being practised.
Ericsson and Harwell, writing from the opposite side of a long argument about practice, concede the same point in Frontiers in Psychology 10 (2019): performance plateaus exist and require adjusted training to overcome. When two camps who agree on almost nothing agree on this, it is worth taking seriously.
The dip that comes before the jump
The most useful detail in the Donner and Hardy paper is the shape of a transition between segments: typically "a sharp decrease followed by a slightly slower increase to a higher level than before the transition." A dip, then a higher plateau.
My log has that shape twice over, and I did not recognise it at the time:
- 96 on 29 July, 40 on 30 July, 94 on 1 August. A 56-mark collapse and a 54-mark recovery in four days, with nothing about my preparation changing in between.
- QUANT falling from 26.8 to 23.6 to 23.4 across three blocks, then jumping to 35.6. The decline ran for ten weeks before the jump.
I want to be careful here, because this is exactly where a post like this turns into astrology. The research describes an average shape across 25,280 curves; it does not license me to tell you your bad mock is a leading indicator. Sometimes a 40 is a 40. What the finding does justify is a decision rule: do not restructure your preparation on the strength of one bad mock, because a dip immediately before an improvement is a documented pattern rather than a rare one. I wrote about not overreacting to a single attempt in the mock analysis post.
Why "stuck at 90 percentile" is partly a measurement problem
If your complaint is specifically that you are stuck around the 90th percentile, some of what you are seeing is the instrument, not you. Percentile is bounded, so the same improvement in marks buys fewer percentile points the higher you go - and mock percentile is computed against whoever sat that particular paper, which varies.
From my own 45 sectional score-to-percentile observations across 15 papers, here are two pairs that should not be possible if percentile were a clean function of score:
| Section | Paper | Score | Percentile |
|---|---|---|---|
| QUANT | SIMCAT 103 | 24 | 91 |
| QUANT | PreSimCAT 03 | 30 | 74 |
| VARC | AIMCAT SA2701 | 25 | 86 |
| VARC | SIMCAT 105 | 39 | 78 |
Six more marks in Quant, seventeen percentile points lower. Fourteen more marks in VARC, eight points lower. Both of those are real rows from real attempts. So a run of 88, 90, 89, 91 across four different papers is not evidence of a ceiling; it may be four different cohorts and four different difficulty levels. The percentile post has all 45 pairs and the spread.
The practical consequence: judge a plateau on section marks across at least five attempts, not on overall percentile across three. Percentile is the better number for knowing where you stand and the worse number for detecting whether you are moving.
What actually breaks a plateau
Four changes, in the order I would try them. None of them is "take more mocks".
- Decompose before you diagnose. Compute your section means in blocks of five attempts, exactly as in the first table. If two sections are moving in opposite directions, you do not have a plateau, you have a neglected section. Mine was DILR, losing 4.2 marks while I congratulated myself about VARC.
- Change the shape of practice, not the volume. If a plateau is the tail of one strategy, the intervention is a different strategy. The best-evidenced version of this for Quant is mixed rather than topic-wise practice: a preregistered trial of 787 students found the mixed group scoring 61% against 38% on a delayed test. The interleaving post has the full trial and its limits.
- Vary deliberately, early in a phase. In Stafford and Dewar's study of 854,064 online game players, higher variance in performance across a player's first five attempts predicted higher scores on attempts six to ten (Pearson r = 0.59, p < 0.0001, with bootstrapped confidence intervals of 0.009 to -0.009 ruling out an artefact of the score distribution). They read this as exploration paying off against exploitation. The transfer to CAT is a guess, but the guess is cheap: when you are flat, trying a genuinely different approach to DILR set selection costs you one mock and might cost you nothing.
- Space the mocks out. In the same dataset, a 24-hour gap between practice sessions was worth roughly 50% extra practice at the same volume. In my own log, seven of nineteen inter-mock gaps were two days or less, and those produced no usable signal at all - there was no room to revise anything in between. Taking more mocks during a plateau usually shortens the gaps, which makes the plateau worse rather than better.
How to tell a plateau from noise
Mock scores are noisy enough that most "plateaus" people report are three attempts long, which is not a plateau. A rule I would defend:
| What you are seeing | What it probably is | What to do |
|---|---|---|
| One bad mock after a good one | Noise, or a documented pre-improvement dip | Analyse it. Change nothing structural |
| Three flat attempts | Too short to read | Keep the schedule. Do not restructure |
| Two blocks of five flat, sections also flat | A real plateau | Change what you practise, not how much |
| Two blocks flat, sections moving oppositely | A neglected section, not a plateau | Move hours to the falling section |
| Percentile flat, marks rising | Percentile compression, or harder papers | Trust the marks. Keep going |
Where this is weak
- My block-four jump is confounded and I will not claim it as a breakthrough. Three of those five attempts were real CAT 2020 papers rather than coaching mocks - a different instrument, differently scaled. If I exclude them, the coaching-mock mean across attempts 11 to 17 is 63.7, against 61.8 for my first five. On coaching mocks alone, there is no leap at all across seventeen attempts. That is the uncomfortable version and it is the honest one.
- Four blocks of five is a tiny sample. A block mean built from five noisy attempts has a wide interval around it that I have not computed and could not meaningfully compute at n = 5. Treat 61.8 versus 59.8 as "the same" rather than as a decline.
- The plateau and its ending are not separable from everything else that changed. My preparation, the mock providers, my sleep and my job all varied across those three months. A personal log has no control condition, so every causal sentence here is a story consistent with the data rather than a finding from it.
- Donner and Hardy studied brain-training tasks, not exams. Their curves are 500 repetitions of a short cognitive task on one commercial platform, with self-selected users. A CAT mock is a three-hour composite instrument taken a few dozen times. The mechanism - averaging hides structure - transfers as a description; none of their numbers transfer as predictions.
- The exploration finding is from a browser game. An r of 0.59 between early variance and later scores in a rapid perception task is a strong result in its own domain and pure extrapolation in mine. I flag it because it is actionable and cheap to test, not because it has been tested here.
- Both camps in the practice literature have positions to defend, and I am citing the point they happen to agree on. That agreement is evidence of something, but it is not the same as an independent test.
- I have not sat CAT. Every percentile above is a mock percentile computed on the cohort that sat that specific paper, which is far smaller and more self-selected than the real exam's.
Sources
- Donner, Y., & Hardy, J. L. (2015). Piecewise power laws in individual learning curves. Psychonomic Bulletin & Review, 22, 1308-1319. Read in full via PMC. The 25,280 curves, the 90.74% against 85.97% fit and the transition shape are from that text.
- Stafford, T., & Dewar, M. (2014). Tracing the trajectory of skill learning with a very large sample of online game players. Psychological Science, 25(2), 511-518. Read in full via the White Rose postprint. The r = 0.59 exploration result and the spacing figure are from that text.
- Ericsson, K. A., & Harwell, K. W. (2019). Deliberate practice and proposed limits on the effects of practice on the acquisition of expert performance. Frontiers in Psychology, 10, 2396. Read in full via PMC.
- Mock log, 20 attempts with per-section scores, 10 May to 8 August 2026, and 45 sectional score-to-percentile observations across 15 papers - first-party, published in the 20-mock post and the percentile post.
The reason I could see the cancellation at all is that the section scores were stored per attempt rather than only the total. Karma Yogi keeps every section by date, which is what turns "my scores are not improving" into a question with an answer.
End of essay
- Anish Guruvelli