Karma YogiKarma Yogi
Start
The Journal
CAT Preparation

10 September 2026
12 min read

A
Anish Guruvelli
Karma Yogi

Overconfidence in Exam Prep: The 14-Point Blind Spot

Across 310 students, the average first-exam prediction was 85% and the average actual score was 71%. The weakest quarter were out by 32 points. Here is why the gap exists, the one intervention that closed it, and the study that found overconfidence also helps.

You score less than you expect because the feeling of knowing something is generated before you have tried to use it. In a study of 310 students in an introductory biology course at the University of Kentucky, the average score predicted for the first exam was 85% and the average score actually earned was 71% - a 14-point gap, and the lowest-performing quarter of the class were out by 32 points. My own version of that number is written in my mock log: on the paper where I scored 2 marks in DILR, my own analysis note from that week reads "Could have definitely gotten 30 marks." The paper was the same paper either way. What changed was which of the two numbers I believed on the day.

This is the most useful thing in the metacognition literature and it is almost never presented with its second half, which is that the gap shrinks on its own for most people, never shrinks for some, and - in one large study - measurably helped the students who had it.

The gap has a measured size

Jennifer Osterhage, Ellen Usher, Trisha Douin and William Bailey published two linked studies in CBE - Life Sciences Education in 2019. Students predicted their score before each exam; the researchers subtracted the actual score to get what they call a discrepancy score. Study 1 followed 310 students across four exams in one semester.

Measure Exam 1 Exam 2 Exam 3 Exam 4
Miscalibrated by 10 points or more59.7%41.7%38.4%27.5%
Overestimated their score59.7%35.5%26.5%8.1%
Underestimated their score0.0%6.2%11.8%19.4%

Three details in that table deserve more attention than the headline.

On the first exam, not one student in 310 underestimated their score by ten points or more. The error is not symmetric noise around the truth. It has a direction, and the direction is up.

By the fourth exam the class as a whole was underestimating, by an average of 2.2%. Four rounds of being wrong in one direction overshot the correction. If your first three mocks have humbled you, do not treat your current pessimism as accuracy either.

The error is worst where the consequences are worst. The lowest-performing quartile overestimated their first-exam score by an average of 32 points. By the final exam that had fallen to 6 points, partly because they predicted better and partly because they scored better. But 5% of the class stayed miscalibrated at every single time point - and those students were among the lowest performing in the class. Calibration is not guaranteed to arrive with experience. It arrives for most people.

The intervention that closed it, in one semester

The second study is the part worth acting on. The following semester, the same instructor taught one section with an added intervention: students were explicitly told about the human tendency to overestimate their own abilities, and given repeated retrieval-practice opportunities (n = 290). A control section of 251 students, taught by a different instructor in the same semester, sat the same first exam.

Exam 1 outcome Study 1 (no intervention) Control section Intervention section
Students miscalibrated by 10+ points59.7%58.7%36.6%
Students who overestimated59.7%53.9%28.0%
Mean discrepancy, same instructor13.5%-3.5%
Exam average71%70%78%
Students scoring below 50%9.5%8.5%3.7%
Lowest quartile overestimate32 points-18 points

Note what the authors say about the mechanism, because it is not what you would guess:

"The improved calibration cannot be explained by a difference in the slope of the predicted lines. The difference in calibration, therefore, is mainly due to improved performance on the first exam by students who received the instructional intervention."

The intervention did not mostly make students predict lower. It made them score higher, by seven to eight points, and the gap closed from the other side. That is the argument for practice testing stated as sharply as I have seen it anywhere: the same activity that tells you what you do not know is the activity that teaches it.

This is also the cleanest available answer to why mock analysis beats mock accumulation. A mock is a calibration instrument and a learning event at the same time, and the second effect was the larger one here.

The result that argues the other way

Jan Magnus and Anatoly Peresetsky published a study in Frontiers in Psychology in 2018 using 592 second-year students at the International College of Economics and Finance in Moscow, across five cohorts from 2011 to 2015. Each student forecast their grade on each of three statistics exams, halfway through the exam itself, with a bonus point for a forecast within 3 marks - which is why the response rate was 97%.

They separated the part of a forecast explained by prior academic results from the part that was not, and called the residual confidence. Then they asked what that residual predicts.

Finding Exam 1 Exam 2 Exam 3
Mean forecast, all years38.2140.9840.06
Mean grade, all years35.3036.9040.27
Effect of unexplained confidence on the grade+0.225+0.282+0.345
How much less overconfident women were (grade points)-5.04-2.99-2.70

All three confidence coefficients are significant at the 1% level. Given the same prior grades, the same homework record and the same year, the more confident student scored higher. Rationality was firmly rejected in all three exams (p-values under 0.2%), and the students were overconfident - and it helped them anyway.

I am including this because leaving it out would let me tell a tidier story than the evidence supports. But there is a second result in the same paper that a CAT aspirant should read very carefully. Doing well on exam 1 made students overpredict exams 2 and 3 (coefficients +0.121 and +0.141, both significant). Doing well on exam 2 made them more cautious about exam 3 (coefficient -0.162). Success early produces overconfidence; success later produces caution. The dangerous mock is the one after your best one.

Mine is public: 96 on 29 July, my best paper to that point, then 40 the next day with 2 marks in DILR. The whole log is here, and that pair of rows is the single most instructive thing in it.

What my own log shows, and what it cannot

I did not record a numeric prediction before each mock, which means I cannot compute a discrepancy score for myself and will not pretend otherwise. What I did record, per section, was a "could have gotten X" figure during analysis - what the paper was worth with better selection.

That number is not a prediction; it is a post-hoc estimate, and it is subject to exactly the bias this post is about. On the 40-mark paper it said 30 marks were available in DILR against the 2 I scored. Either the paper was genuinely worth 28 more marks to me, in which case my problem is entirely triage, or the estimate is inflated by hindsight, in which case my problem includes not knowing what I can solve. I cannot separate those two from my own data, and I do not think most people writing "could have gotten X" in a mock analysis can either.

What my log does contain is the volatility that makes any single self-assessment unreliable: across 19 consecutive mock-to-mock transitions, my overall score moved by an average of 19.9 marks, with a median of 15 and a maximum of 56. If your sense of how prepared you are updates after every mock, it is updating on something that swings 20 marks for reasons that are not your preparation. That is the subject of a separate post.

What to actually do about it

  • Write the number down before the mock, not after. A prediction you did not record is a prediction you will retro-fit. Two minutes, three section numbers, before you open the paper.
  • Expect the first few to be badly wrong and in one direction. 59.7% of that biology class overestimated on exam 1 and none underestimated. Being 15 marks out on your first three mocks is the normal case, not a character flaw.
  • Do not over-correct. By exam 4 the same class was underestimating by 2.2% on average. A pessimistic prediction is a miscalibrated prediction too, and it produces the opposite bad decision - abandoning sets you could have solved.
  • Solve, do not re-read. The intervention that halved miscalibration paired an explicit warning about overestimation with retrieval practice, and its effect ran mostly through better scores rather than lower predictions. Re-reading a solution generates the feeling of knowing without testing it, which is the machinery that makes the gap in the first place. There is separate evidence on that specific comparison.
  • Watch the mock after a good mock. Success on an early assessment made those Moscow students overpredict the next one. My 96 was followed by a 40.
  • If your predictions never converge, that is the signal. 5% of the biology class stayed miscalibrated at every single point, and they were among the lowest performers. Persistent miscalibration is diagnostic, not cosmetic.

Where this evidence is weak

  • Neither study is about CAT, and one is not even about a hard exam. Introductory biology at a US public university, graded in percentages, with four exams a semester. CAT is one three-hour paper with sectional locks and negative marking, sat by roughly 2.5 lakh graduates. The direction of the bias is likely universal. The 14-point figure is not a prediction about your mock.
  • The intervention study is not a randomised trial. The comparison sections were taught by different instructors in different semesters, and the authors say plainly that they "cannot rule out the possibility that demographic differences between sections of the course and between semesters may have affected our results". They also combined two interventions, so nobody knows which half did the work.
  • Withdrawal rates differed between the sections - 4.92% in the intervention section against 13.6% in the control - and only students who sat every exam were analysed. Losing more of the weakest students from one group is exactly the kind of thing that flatters a comparison.
  • The Moscow forecasts were made halfway through the exam. Students had already seen and answered part one, which makes their forecast far better informed than a prediction made the week before a mock. Whether "overconfidence helps" survives when the forecast is made in advance is not established by that paper.
  • Overconfidence helping is a correlation with a plausible reverse story. Confident students may work harder, or may simply have private information the model does not see - the authors list tutoring and psychological traits as unmeasured. They do not claim causation and neither do I.
  • My own log has no recorded predictions, so the only first-party number here is a post-hoc "could have gotten" estimate that is itself vulnerable to hindsight bias, plus one candidate, one quarter, and no CAT score to check anything against.
  • Nobody has measured calibration in Indian competitive exam prep. Not for CAT. If somebody has, I could not find it, and the aspirant population - graduates, self-selected, sitting 15 to 20 mocks - is different enough from a first-semester biology class that I would not assume the numbers carry.

Sources

  • Osterhage, J. L., Usher, E. L., Douin, T. A., & Bailey, W. M. (2019). Opportunities for self-evaluation increase student calibration in an introductory biology course. CBE - Life Sciences Education, 18(2):ar16. Open access at PubMed Central. All calibration figures and both tables above are from its Results and Tables 1-3.
  • Magnus, J. R., & Peresetsky, A. A. (2018). Grade expectations: rationality and overconfidence. Frontiers in Psychology, 8:2346. Open access at PubMed Central. Coefficients from its Tables 1, 2, 4 and 5.
  • Mock log, 20 attempts, 10 May to 8 August 2026, including the 96 and the 40 - first-party, published in full in the 20-mock post. The 19.9-mark mean absolute change is computed from those 20 rows.

The cheapest version of all of this is a number written down before the paper and compared after it. Karma Yogi keeps section scores and analysis notes against each mock, so a prediction logged in the notes is still there when the result arrives - which is the only way the comparison survives your memory of what you expected.

End of essay

- Anish Guruvelli

Common questions

Why do I score less in mocks than I expect to?
Because the sense that you know something is produced by recognising it, not by retrieving it, and a mock is the first moment those two come apart. Across 310 students in one course, the average predicted first-exam score was 85% against an actual 71%, and not a single student underestimated by ten points or more.
What is a judgment of learning?
It is your own estimate of how well you have learnt something, made before you are tested on it. The research interest is in how badly it tracks reality: people routinely rate familiar material as learnt when they cannot yet produce it, which is why a practice test and a re-read feel similar and score differently.
Does overconfidence go away with more mocks?
For most people, gradually. In that biology course the share of students miscalibrated by ten points or more fell from 59.7% on exam 1 to 27.5% by exam 4. But 5% of the class was still miscalibrated at every measurement, and those students were among the lowest performers, so it is not automatic.
Is being confident bad for exam performance?
Not obviously. A study of 592 students in Moscow separated the part of a grade forecast that prior results could not explain and found that this residual confidence positively predicted the actual grade in all three exams, significant at the 1% level. Miscalibration and confidence are different things, and only one of them is clearly costly.
How do I know if I am overestimating my preparation?
Write a predicted score for each section before you open the mock, then compare. If you have never written the number down you will reconstruct it after seeing the result, which is the one procedure guaranteed to make you look calibrated.
Why does re-reading make me feel prepared when I am not?
Re-reading generates fluency, and fluency feels like knowledge. You recognise the material easily, so you rate yourself as knowing it, and nothing in the activity ever asks you to produce it from nothing. That is the exact condition under which self-assessments are least accurate.
Is the Dunning-Kruger effect real for exam scores?
In this dataset, yes, in the narrow sense that error size tracked performance. The lowest-performing quartile overestimated their first exam by 32 points while the highest quartile were close or slightly under. Whether the broader interpretation of the effect holds up is a separate and much more contested argument.
What actually fixed students overconfidence in the study?
Being told explicitly that people overestimate themselves, plus repeated retrieval practice. Miscalibration in that section fell from 58.7% to 36.6%. The interesting part is the mechanism: the authors found the gap closed mainly because those students scored eight points higher, not because they predicted lower.
Should I predict my score before every mock?
Yes, and it takes two minutes. Three section numbers written down before you open the paper, compared afterwards. A prediction you keep in your head is not a measurement, because you will adjust it retroactively once you know the answer.
Why did I do badly right after my best mock?
Partly ordinary score volatility, and partly this. In the Moscow data, doing well on the first exam made students significantly overpredict the second one, while doing well on the second made them more cautious about the third. My own log has a 96 followed the next day by a 40 with 2 marks in DILR.
Can you be underconfident, and does that hurt too?
Yes to both. By the fourth exam that biology class had overshot the correction and was underestimating scores by an average of 2.2%. In an exam with set selection under a clock, underconfidence costs you attempts you would have got right, which is a different failure with the same root cause.
Does any of this research use Indian exam data?
No. The calibration work here is a US biology course and a Moscow economics programme, and I could not find a study measuring prediction accuracy in Indian competitive exam preparation. Treat the direction of the bias as likely universal and the specific numbers as not transferable.
Keep Reading

More on CAT preparation.

All CAT guides
CAT Preparation· 12 min read

Deliberate Practice for CAT: What 88 Studies Actually Say

The "quality beats quantity" claim has a real evidence base and a much smaller effect than the popular version admits: 12% of performance variance overall, 4% in education. Here is what survives contact with the meta-analysis, and what my own DILR log says about practising a section for three months and moving two marks.

Read
Study Habits· 12 min read

How to Build a Study Habit: 66 Days Is a Median, Not a Law

The "66 days to form a habit" number comes from 39 people drinking water and doing sit-ups, and the actual range was 18 to 254 days. Here is what that study really found, what it says about a study streak, and why the same paper is the best argument for not punishing a missed day.

Read
Study Techniques· 12 min read

Best Time to Study Morning or Night: What 3.4M Logins Say

There is no universal best hour, and the largest datasets on this question disagree about which direction the effect even runs. Here is what four studies actually found, including the one whose own prediction was refuted, and a three-week test you can run on your own log instead.

Read
CAT Preparation· 13 min read

How Many Hours a Day Should You Study for CAT? I Logged 335

My weekly goal was 1,260 minutes and my average active day is 149 minutes, which needs 8.5 days a week to hit. Here is the arithmetic that broke my own target, what 335 logged hours actually bought, and why the hours literature is far shakier than the 10,000-hour story suggests.

Read