Karma YogiKarma Yogi
Start
The Journal
CAT Mocks

10 September 2026
13 min read

A
Anish Guruvelli
Karma Yogi

Why Mock Scores Plateau: 10 Weeks Flat in My Own Log

My overall mock score averaged 61.8, then 59.8, then 62.0 across three consecutive blocks of five - ten weeks of nothing. Underneath it VARC was gaining 7.8 marks and Quant was losing 3.4. The plateau was two real trends cancelling, which is also what the learning-curve research says a plateau usually is.

A plateau in your overall mock score is almost never a plateau in your ability - it is usually two real trends cancelling, and the fix is a change in what you practise rather than more of it. In my own log of 20 CAT mocks, the mean overall score across three consecutive blocks of five went 61.8, then 59.8, then 62.0. Ten weeks, no movement. Underneath that flat line, my VARC mean rose from 16.6 to 24.4 while my QUANT mean fell from 26.8 to 23.4. Two genuine trends, pointed in opposite directions, summing to nothing. The research says the same thing at a much larger scale: an analysis of 25,280 individual learning curves found that the smooth improvement curve everybody quotes is largely an artefact of averaging, and that real individual curves contain plateaus, dips and sudden jumps.

So "my mock scores are not improving" is usually the wrong diagnosis. The right question is which of your three sections is moving, in which direction, and what you would have to change to stop them cancelling.

My actual plateau, in blocks of five

Twenty full mocks between 10 May and 8 August 2026, grouped in consecutive blocks of five, scores out of 198. The last three attempts in block four are real CAT 2020 papers rather than coaching mocks - a confound I deal with properly further down rather than burying.

Block Overall (mean) Percentile (mean) VARC DILR QUANT
Mocks 1-561.873.416.618.426.8
Mocks 6-1059.880.217.019.223.6
Mocks 11-1562.080.524.414.223.4
Mocks 16-2081.094.324.820.635.6

Read the middle two rows. Ten weeks and ten mocks in which the number I was actually looking at - the overall score - moved by 2.2 marks in the wrong direction and then 2.2 back. If you had asked me in mid-July whether I was improving, I would have said no, and I would have been wrong in a specific and useful way.

The plateau was three sections cancelling

Between block one and block three, with the overall flat:

  • VARC: 16.6 to 24.4. Up 7.8 marks. This was the daily reading habit and the RC question-type work finally landing.
  • QUANT: 26.8 to 23.4. Down 3.4 marks. My strongest section going backwards while I paid it no attention, precisely because it was my strongest.
  • DILR: 18.4 to 14.2. Down 4.2 marks, and this was the real problem hiding under the flat line.

Add those up and you get roughly zero, which is exactly what the overall column shows. An aggregate that is not moving is not evidence that nothing is moving. It is a single number doing what single numbers do - averaging away the information you need.

This is the same failure mode, one level down, that the learning-curve literature describes at population scale, and I find that symmetry genuinely useful: whether you average across three sections or across twenty-two thousand people, averaging manufactures smoothness that is not in the underlying data.

What the learning-curve research actually says about plateaus

Donner and Hardy, Psychonomic Bulletin & Review 22:1308-1319, 2015. They analysed 25,280 learning curves from 22,460 people on the Lumosity cognitive training platform, each curve carrying 500 measurements of performance on the same task. That is an unusually good look at what individual learning actually looks like, because almost every classic study of practice averaged across participants first and drew a curve second.

Three findings that change how you should read your own mock log:

  • Individual curves are not smooth. A piecewise power law - a curve made of discrete segments with breakpoints - explained 90.74% of variance against 85.97% for a single smooth power law, a difference so large it is hard to overstate statistically (chi-squared (94205) = 534,030). Two- and three-segment solutions were the most common.
  • Real curves contain "both plateaus and bursts of rapid improvement". The smooth curve, the authors write, "may be an artifact of averaging across individual curves". The tidy diminishing-returns picture you have seen is a population statistic, not a description of anybody.
  • The segments are strategy changes. They describe the pattern as "a discrete sequence of strategy shifts, in which each strategy is better in the long term than the ones preceding it".

That third point is the practically important one and it is the whole argument of this post. If a plateau is the tail of one strategy rather than a limit on your ability, then more repetitions of that strategy is the one intervention guaranteed not to work. You do not out-practise a plateau. You change what is being practised.

Ericsson and Harwell, writing from the opposite side of a long argument about practice, concede the same point in Frontiers in Psychology 10 (2019): performance plateaus exist and require adjusted training to overcome. When two camps who agree on almost nothing agree on this, it is worth taking seriously.

The dip that comes before the jump

The most useful detail in the Donner and Hardy paper is the shape of a transition between segments: typically "a sharp decrease followed by a slightly slower increase to a higher level than before the transition." A dip, then a higher plateau.

My log has that shape twice over, and I did not recognise it at the time:

  • 96 on 29 July, 40 on 30 July, 94 on 1 August. A 56-mark collapse and a 54-mark recovery in four days, with nothing about my preparation changing in between.
  • QUANT falling from 26.8 to 23.6 to 23.4 across three blocks, then jumping to 35.6. The decline ran for ten weeks before the jump.

I want to be careful here, because this is exactly where a post like this turns into astrology. The research describes an average shape across 25,280 curves; it does not license me to tell you your bad mock is a leading indicator. Sometimes a 40 is a 40. What the finding does justify is a decision rule: do not restructure your preparation on the strength of one bad mock, because a dip immediately before an improvement is a documented pattern rather than a rare one. I wrote about not overreacting to a single attempt in the mock analysis post.

Why "stuck at 90 percentile" is partly a measurement problem

If your complaint is specifically that you are stuck around the 90th percentile, some of what you are seeing is the instrument, not you. Percentile is bounded, so the same improvement in marks buys fewer percentile points the higher you go - and mock percentile is computed against whoever sat that particular paper, which varies.

From my own 45 sectional score-to-percentile observations across 15 papers, here are two pairs that should not be possible if percentile were a clean function of score:

SectionPaperScorePercentile
QUANTSIMCAT 1032491
QUANTPreSimCAT 033074
VARCAIMCAT SA27012586
VARCSIMCAT 1053978

Six more marks in Quant, seventeen percentile points lower. Fourteen more marks in VARC, eight points lower. Both of those are real rows from real attempts. So a run of 88, 90, 89, 91 across four different papers is not evidence of a ceiling; it may be four different cohorts and four different difficulty levels. The percentile post has all 45 pairs and the spread.

The practical consequence: judge a plateau on section marks across at least five attempts, not on overall percentile across three. Percentile is the better number for knowing where you stand and the worse number for detecting whether you are moving.

What actually breaks a plateau

Four changes, in the order I would try them. None of them is "take more mocks".

  1. Decompose before you diagnose. Compute your section means in blocks of five attempts, exactly as in the first table. If two sections are moving in opposite directions, you do not have a plateau, you have a neglected section. Mine was DILR, losing 4.2 marks while I congratulated myself about VARC.
  2. Change the shape of practice, not the volume. If a plateau is the tail of one strategy, the intervention is a different strategy. The best-evidenced version of this for Quant is mixed rather than topic-wise practice: a preregistered trial of 787 students found the mixed group scoring 61% against 38% on a delayed test. The interleaving post has the full trial and its limits.
  3. Vary deliberately, early in a phase. In Stafford and Dewar's study of 854,064 online game players, higher variance in performance across a player's first five attempts predicted higher scores on attempts six to ten (Pearson r = 0.59, p < 0.0001, with bootstrapped confidence intervals of 0.009 to -0.009 ruling out an artefact of the score distribution). They read this as exploration paying off against exploitation. The transfer to CAT is a guess, but the guess is cheap: when you are flat, trying a genuinely different approach to DILR set selection costs you one mock and might cost you nothing.
  4. Space the mocks out. In the same dataset, a 24-hour gap between practice sessions was worth roughly 50% extra practice at the same volume. In my own log, seven of nineteen inter-mock gaps were two days or less, and those produced no usable signal at all - there was no room to revise anything in between. Taking more mocks during a plateau usually shortens the gaps, which makes the plateau worse rather than better.

How to tell a plateau from noise

Mock scores are noisy enough that most "plateaus" people report are three attempts long, which is not a plateau. A rule I would defend:

What you are seeingWhat it probably isWhat to do
One bad mock after a good oneNoise, or a documented pre-improvement dipAnalyse it. Change nothing structural
Three flat attemptsToo short to readKeep the schedule. Do not restructure
Two blocks of five flat, sections also flatA real plateauChange what you practise, not how much
Two blocks flat, sections moving oppositelyA neglected section, not a plateauMove hours to the falling section
Percentile flat, marks risingPercentile compression, or harder papersTrust the marks. Keep going

Where this is weak

  • My block-four jump is confounded and I will not claim it as a breakthrough. Three of those five attempts were real CAT 2020 papers rather than coaching mocks - a different instrument, differently scaled. If I exclude them, the coaching-mock mean across attempts 11 to 17 is 63.7, against 61.8 for my first five. On coaching mocks alone, there is no leap at all across seventeen attempts. That is the uncomfortable version and it is the honest one.
  • Four blocks of five is a tiny sample. A block mean built from five noisy attempts has a wide interval around it that I have not computed and could not meaningfully compute at n = 5. Treat 61.8 versus 59.8 as "the same" rather than as a decline.
  • The plateau and its ending are not separable from everything else that changed. My preparation, the mock providers, my sleep and my job all varied across those three months. A personal log has no control condition, so every causal sentence here is a story consistent with the data rather than a finding from it.
  • Donner and Hardy studied brain-training tasks, not exams. Their curves are 500 repetitions of a short cognitive task on one commercial platform, with self-selected users. A CAT mock is a three-hour composite instrument taken a few dozen times. The mechanism - averaging hides structure - transfers as a description; none of their numbers transfer as predictions.
  • The exploration finding is from a browser game. An r of 0.59 between early variance and later scores in a rapid perception task is a strong result in its own domain and pure extrapolation in mine. I flag it because it is actionable and cheap to test, not because it has been tested here.
  • Both camps in the practice literature have positions to defend, and I am citing the point they happen to agree on. That agreement is evidence of something, but it is not the same as an independent test.
  • I have not sat CAT. Every percentile above is a mock percentile computed on the cohort that sat that specific paper, which is far smaller and more self-selected than the real exam's.

Sources

  • Donner, Y., & Hardy, J. L. (2015). Piecewise power laws in individual learning curves. Psychonomic Bulletin & Review, 22, 1308-1319. Read in full via PMC. The 25,280 curves, the 90.74% against 85.97% fit and the transition shape are from that text.
  • Stafford, T., & Dewar, M. (2014). Tracing the trajectory of skill learning with a very large sample of online game players. Psychological Science, 25(2), 511-518. Read in full via the White Rose postprint. The r = 0.59 exploration result and the spacing figure are from that text.
  • Ericsson, K. A., & Harwell, K. W. (2019). Deliberate practice and proposed limits on the effects of practice on the acquisition of expert performance. Frontiers in Psychology, 10, 2396. Read in full via PMC.
  • Mock log, 20 attempts with per-section scores, 10 May to 8 August 2026, and 45 sectional score-to-percentile observations across 15 papers - first-party, published in the 20-mock post and the percentile post.

The reason I could see the cancellation at all is that the section scores were stored per attempt rather than only the total. Karma Yogi keeps every section by date, which is what turns "my scores are not improving" into a question with an answer.

End of essay

- Anish Guruvelli

Common questions

Why are my mock scores not improving?
Most often because two sections are moving in opposite directions and cancelling in the total. In my own log the overall mean was flat at 61.8, 59.8 and 62.0 across ten weeks while VARC gained 7.8 marks and Quant lost 3.4. Break the total into sections before concluding anything.
How long does a mock score plateau usually last?
There is no established figure for CAT and I would distrust anyone giving one. Mine ran roughly ten weeks and ten attempts. What the learning-curve research does say is that plateaus are a normal segment of an individual curve rather than a sign something has gone wrong.
Is a plateau a sign I have hit my ceiling?
Very unlikely. An analysis of 25,280 individual learning curves found the typical curve is made of discrete segments described as a sequence of strategy shifts, each better than the last. A plateau is usually the tail of one strategy, not a limit on the person using it.
Should I take more mocks to break a plateau?
Usually the opposite. Taking more mocks compresses the gaps between them, and in my log the seven gaps of two days or less produced no usable signal because nothing could be revised in between. Change what you practise rather than how often you test it.
I am stuck at 90 percentile. What do I do?
First check whether it is real. Percentile compresses near the top and is computed against whoever sat that paper. In my own data a Quant score of 24 returned the 91st percentile on one paper while a 30 returned the 74th on another. Judge movement on section marks, not percentile.
Does a bad mock mean something is wrong?
Not necessarily, and there is a documented pattern of a sharp dip immediately before a step up in individual learning curves. Mine went 96, then 40, then 94 in four days with nothing changing. Analyse the bad attempt; do not restructure your plan on the strength of it.
How many mocks do I need before I can call it a plateau?
Two blocks of five, with the section scores also flat. Three flat attempts is too short to read given how much papers vary. If the sections are moving in opposite directions across those ten attempts, that is a neglected section rather than a plateau.
Which section usually causes a hidden plateau?
In my log it was DILR, which fell 4.2 marks across the flat stretch while VARC rose and masked it. DILR responds to exposure across set types over months rather than to topic drilling, so it degrades quietly when attention moves elsewhere.
Can changing my practice really break a plateau?
It is the intervention the evidence points at. A preregistered trial of 787 students found mixed practice beating topic-wise practice 61% to 38% on a delayed test, and both sides of the long argument about deliberate practice agree that plateaus require adjusted training rather than more of the same.
Why do averaged learning curves look smooth if real ones are not?
Because averaging manufactures smoothness. Donner and Hardy write that the classic smooth power law may be an artifact of averaging across individual curves; a piecewise model fitted their 25,280 curves substantially better. The same thing happens one level down when three sections are averaged into one total.
Should I trust my raw score or my percentile during a plateau?
Section marks for detecting movement, percentile for knowing where you stand. Marks are comparable within a section across papers of similar design; percentile depends on the cohort that sat that particular paper, which changes every time.