You score less than you expect because the feeling of knowing something is generated before you have tried to use it. In a study of 310 students in an introductory biology course at the University of Kentucky, the average score predicted for the first exam was 85% and the average score actually earned was 71% - a 14-point gap, and the lowest-performing quarter of the class were out by 32 points. My own version of that number is written in my mock log: on the paper where I scored 2 marks in DILR, my own analysis note from that week reads "Could have definitely gotten 30 marks." The paper was the same paper either way. What changed was which of the two numbers I believed on the day.
This is the most useful thing in the metacognition literature and it is almost never presented with its second half, which is that the gap shrinks on its own for most people, never shrinks for some, and - in one large study - measurably helped the students who had it.
The gap has a measured size
Jennifer Osterhage, Ellen Usher, Trisha Douin and William Bailey published two linked studies in CBE - Life Sciences Education in 2019. Students predicted their score before each exam; the researchers subtracted the actual score to get what they call a discrepancy score. Study 1 followed 310 students across four exams in one semester.
| Measure | Exam 1 | Exam 2 | Exam 3 | Exam 4 |
|---|---|---|---|---|
| Miscalibrated by 10 points or more | 59.7% | 41.7% | 38.4% | 27.5% |
| Overestimated their score | 59.7% | 35.5% | 26.5% | 8.1% |
| Underestimated their score | 0.0% | 6.2% | 11.8% | 19.4% |
Three details in that table deserve more attention than the headline.
On the first exam, not one student in 310 underestimated their score by ten points or more. The error is not symmetric noise around the truth. It has a direction, and the direction is up.
By the fourth exam the class as a whole was underestimating, by an average of 2.2%. Four rounds of being wrong in one direction overshot the correction. If your first three mocks have humbled you, do not treat your current pessimism as accuracy either.
The error is worst where the consequences are worst. The lowest-performing quartile overestimated their first-exam score by an average of 32 points. By the final exam that had fallen to 6 points, partly because they predicted better and partly because they scored better. But 5% of the class stayed miscalibrated at every single time point - and those students were among the lowest performing in the class. Calibration is not guaranteed to arrive with experience. It arrives for most people.
The intervention that closed it, in one semester
The second study is the part worth acting on. The following semester, the same instructor taught one section with an added intervention: students were explicitly told about the human tendency to overestimate their own abilities, and given repeated retrieval-practice opportunities (n = 290). A control section of 251 students, taught by a different instructor in the same semester, sat the same first exam.
| Exam 1 outcome | Study 1 (no intervention) | Control section | Intervention section |
|---|---|---|---|
| Students miscalibrated by 10+ points | 59.7% | 58.7% | 36.6% |
| Students who overestimated | 59.7% | 53.9% | 28.0% |
| Mean discrepancy, same instructor | 13.5% | - | 3.5% |
| Exam average | 71% | 70% | 78% |
| Students scoring below 50% | 9.5% | 8.5% | 3.7% |
| Lowest quartile overestimate | 32 points | - | 18 points |
Note what the authors say about the mechanism, because it is not what you would guess:
"The improved calibration cannot be explained by a difference in the slope of the predicted lines. The difference in calibration, therefore, is mainly due to improved performance on the first exam by students who received the instructional intervention."
The intervention did not mostly make students predict lower. It made them score higher, by seven to eight points, and the gap closed from the other side. That is the argument for practice testing stated as sharply as I have seen it anywhere: the same activity that tells you what you do not know is the activity that teaches it.
This is also the cleanest available answer to why mock analysis beats mock accumulation. A mock is a calibration instrument and a learning event at the same time, and the second effect was the larger one here.
The result that argues the other way
Jan Magnus and Anatoly Peresetsky published a study in Frontiers in Psychology in 2018 using 592 second-year students at the International College of Economics and Finance in Moscow, across five cohorts from 2011 to 2015. Each student forecast their grade on each of three statistics exams, halfway through the exam itself, with a bonus point for a forecast within 3 marks - which is why the response rate was 97%.
They separated the part of a forecast explained by prior academic results from the part that was not, and called the residual confidence. Then they asked what that residual predicts.
| Finding | Exam 1 | Exam 2 | Exam 3 |
|---|---|---|---|
| Mean forecast, all years | 38.21 | 40.98 | 40.06 |
| Mean grade, all years | 35.30 | 36.90 | 40.27 |
| Effect of unexplained confidence on the grade | +0.225 | +0.282 | +0.345 |
| How much less overconfident women were (grade points) | -5.04 | -2.99 | -2.70 |
All three confidence coefficients are significant at the 1% level. Given the same prior grades, the same homework record and the same year, the more confident student scored higher. Rationality was firmly rejected in all three exams (p-values under 0.2%), and the students were overconfident - and it helped them anyway.
I am including this because leaving it out would let me tell a tidier story than the evidence supports. But there is a second result in the same paper that a CAT aspirant should read very carefully. Doing well on exam 1 made students overpredict exams 2 and 3 (coefficients +0.121 and +0.141, both significant). Doing well on exam 2 made them more cautious about exam 3 (coefficient -0.162). Success early produces overconfidence; success later produces caution. The dangerous mock is the one after your best one.
Mine is public: 96 on 29 July, my best paper to that point, then 40 the next day with 2 marks in DILR. The whole log is here, and that pair of rows is the single most instructive thing in it.
What my own log shows, and what it cannot
I did not record a numeric prediction before each mock, which means I cannot compute a discrepancy score for myself and will not pretend otherwise. What I did record, per section, was a "could have gotten X" figure during analysis - what the paper was worth with better selection.
That number is not a prediction; it is a post-hoc estimate, and it is subject to exactly the bias this post is about. On the 40-mark paper it said 30 marks were available in DILR against the 2 I scored. Either the paper was genuinely worth 28 more marks to me, in which case my problem is entirely triage, or the estimate is inflated by hindsight, in which case my problem includes not knowing what I can solve. I cannot separate those two from my own data, and I do not think most people writing "could have gotten X" in a mock analysis can either.
What my log does contain is the volatility that makes any single self-assessment unreliable: across 19 consecutive mock-to-mock transitions, my overall score moved by an average of 19.9 marks, with a median of 15 and a maximum of 56. If your sense of how prepared you are updates after every mock, it is updating on something that swings 20 marks for reasons that are not your preparation. That is the subject of a separate post.
What to actually do about it
- Write the number down before the mock, not after. A prediction you did not record is a prediction you will retro-fit. Two minutes, three section numbers, before you open the paper.
- Expect the first few to be badly wrong and in one direction. 59.7% of that biology class overestimated on exam 1 and none underestimated. Being 15 marks out on your first three mocks is the normal case, not a character flaw.
- Do not over-correct. By exam 4 the same class was underestimating by 2.2% on average. A pessimistic prediction is a miscalibrated prediction too, and it produces the opposite bad decision - abandoning sets you could have solved.
- Solve, do not re-read. The intervention that halved miscalibration paired an explicit warning about overestimation with retrieval practice, and its effect ran mostly through better scores rather than lower predictions. Re-reading a solution generates the feeling of knowing without testing it, which is the machinery that makes the gap in the first place. There is separate evidence on that specific comparison.
- Watch the mock after a good mock. Success on an early assessment made those Moscow students overpredict the next one. My 96 was followed by a 40.
- If your predictions never converge, that is the signal. 5% of the biology class stayed miscalibrated at every single point, and they were among the lowest performers. Persistent miscalibration is diagnostic, not cosmetic.
Where this evidence is weak
- Neither study is about CAT, and one is not even about a hard exam. Introductory biology at a US public university, graded in percentages, with four exams a semester. CAT is one three-hour paper with sectional locks and negative marking, sat by roughly 2.5 lakh graduates. The direction of the bias is likely universal. The 14-point figure is not a prediction about your mock.
- The intervention study is not a randomised trial. The comparison sections were taught by different instructors in different semesters, and the authors say plainly that they "cannot rule out the possibility that demographic differences between sections of the course and between semesters may have affected our results". They also combined two interventions, so nobody knows which half did the work.
- Withdrawal rates differed between the sections - 4.92% in the intervention section against 13.6% in the control - and only students who sat every exam were analysed. Losing more of the weakest students from one group is exactly the kind of thing that flatters a comparison.
- The Moscow forecasts were made halfway through the exam. Students had already seen and answered part one, which makes their forecast far better informed than a prediction made the week before a mock. Whether "overconfidence helps" survives when the forecast is made in advance is not established by that paper.
- Overconfidence helping is a correlation with a plausible reverse story. Confident students may work harder, or may simply have private information the model does not see - the authors list tutoring and psychological traits as unmeasured. They do not claim causation and neither do I.
- My own log has no recorded predictions, so the only first-party number here is a post-hoc "could have gotten" estimate that is itself vulnerable to hindsight bias, plus one candidate, one quarter, and no CAT score to check anything against.
- Nobody has measured calibration in Indian competitive exam prep. Not for CAT. If somebody has, I could not find it, and the aspirant population - graduates, self-selected, sitting 15 to 20 mocks - is different enough from a first-semester biology class that I would not assume the numbers carry.
Sources
- Osterhage, J. L., Usher, E. L., Douin, T. A., & Bailey, W. M. (2019). Opportunities for self-evaluation increase student calibration in an introductory biology course. CBE - Life Sciences Education, 18(2):ar16. Open access at PubMed Central. All calibration figures and both tables above are from its Results and Tables 1-3.
- Magnus, J. R., & Peresetsky, A. A. (2018). Grade expectations: rationality and overconfidence. Frontiers in Psychology, 8:2346. Open access at PubMed Central. Coefficients from its Tables 1, 2, 4 and 5.
- Mock log, 20 attempts, 10 May to 8 August 2026, including the 96 and the 40 - first-party, published in full in the 20-mock post. The 19.9-mark mean absolute change is computed from those 20 rows.
The cheapest version of all of this is a number written down before the paper and compared after it. Karma Yogi keeps section scores and analysis notes against each mock, so a prediction logged in the notes is still there when the result arrives - which is the only way the comparison survives your memory of what you expected.
End of essay
- Anish Guruvelli