There is no hour count that buys a CAT percentile, and I can now show that on my own numbers rather than assert it. Between 10 May and 8 August 2026 I logged 20,100 minutes - 335 hours - across 340 sessions, and sat 20 full mocks in the same window. Fit a line through those 20 overall scores against the date and preparation time explains 13.7% of the variation in any single result. Look at the more direct question - do more days of preparation between two mocks produce a better second mock? - and the correlation is r = 0.036. That is zero to three decimal places of usefulness.
The hours still mattered. What they did not do is anything you could see in one mock, which is the only place most aspirants ever look. This post is the whole calculation, including the parts that do not support the tidy answer.
The hours, first
So that the numbers below have something concrete behind them, this is the log they come from.
| Measure | Value |
|---|---|
| Total logged minutes | 20,100 (335 hours) |
| Logged sessions | 340 |
| Distinct active days | 135 |
| Mean session length | 59 minutes |
| Longest single session | 190 minutes |
| Minutes per active day | 149 |
| Weekly goal set | 1,260 minutes (21 hours) |
| Full mocks in the window | 20, over 90 days |
Two honest caveats before anything is derived from this. The session totals are the lifetime figures for the log, so I cannot cleanly attribute every one of the 335 hours to the 90-day mock window - the bulk of it falls there, but I am not going to pretend to a split I do not have. And I was working a full-time job throughout, which is what produces the 59-minute average session; the shape of that week is a separate post.
Question one: does more time between mocks produce a better mock?
This is the cleanest test available in the data. The gap between two consecutive mocks is preparation time you definitely had. If hours convert to marks in any direct way, a 7-day gap should beat a 1-day gap.
| Date | Days since last mock | Change in score | Score |
|---|---|---|---|
| 16 May | 6 | +20 | 61 |
| 23 May | 7 | +13 | 74 |
| 25 May | 2 | -6 | 68 |
| 30 May | 5 | -3 | 65 |
| 6 Jun | 7 | -21 | 44 |
| 13 Jun | 7 | +19 | 63 |
| 20 Jun | 7 | +28 | 91 |
| 27 Jun | 7 | -37 | 54 |
| 28 Jun | 1 | -7 | 47 |
| 5 Jul | 7 | +2 | 49 |
| 11 Jul | 6 | +31 | 80 |
| 18 Jul | 7 | -15 | 65 |
| 25 Jul | 7 | -9 | 56 |
| 27 Jul | 2 | +4 | 60 |
| 29 Jul | 2 | +36 | 96 |
| 30 Jul | 1 | -56 | 40 |
| 1 Aug | 2 | +54 | 94 |
| 3 Aug | 2 | -10 | 84 |
| 8 Aug | 5 | +7 | 91 |
Correlation between the gap and the score change: r = 0.036, r-squared 0.001. Between the gap and the resulting score: r = -0.107, i.e. very slightly negative. Split it crudely and the picture does not improve for the hours hypothesis - the seven mocks that followed a gap of two days or less averaged 69.9, and the twelve that followed a gap of five days or more averaged 66.1.
That comparison is confounded and I am not going to present it as a finding. Six of the seven short-gap mocks fall in the last three weeks of the window, when I was simply better at the exam than I had been in May. The short gaps did not cause the higher scores; they happened to coincide with the end of the period. It is a good illustration of why a clean-looking two-group comparison from n = 19 observations should not persuade anybody of anything, including the person who computed it.
What survives the confound is the correlation of 0.036, which is not vulnerable to it in the same direction: whatever the days between two mocks are doing, they are not predicting how much the score moves.
Question two: does the score trend upward with accumulated preparation?
Yes, and much less than you would hope from one mock.
- Trend: +6.83 marks per 30 days, across the 90-day window
- Correlation with elapsed time: r = 0.370, so r-squared = 0.137
- Residual scatter about that trend: standard deviation 16.9 marks
- Raw spread of the 20 scores: standard deviation 18.2 marks, mean 66.2
Read those together and the conclusion is stark. Accumulated preparation moves my expected score by roughly seven marks a month. The noise around that expectation is seventeen marks. So the signal from a full month of study is comfortably smaller than the difference between a good day and a bad one - which means a single mock result carries almost no information about whether the month worked.
The most extreme demonstration of this is already in the raw log: 96 on 29 July, then 40 on 30 July. A 56-mark swing overnight. My entire three-month improvement, measured as first-five average against last-five average, is 19.2 marks. The one-day swing is 2.9 times the size of the quarter's progress. Nothing about my preparation changed between those two Wednesdays.
Question three: does the relationship reappear if you stop reading single mocks?
This is where the honest answer turns constructive. Take a five-mock rolling average instead of individual results and refit the same trend:
| What is fitted | Trend / 30 days | r-squared | Residual SD |
|---|---|---|---|
| Individual mock scores | +6.83 | 0.137 | 16.9 |
| Five-mock rolling average | +4.20 | 0.264 | 5.3 |
The residual scatter falls from 16.9 marks to 5.3 - a bit over a threefold reduction - and the fit roughly doubles. The rolling average itself runs 61.8 through the end of May and 81.0 by 8 August, which is the improvement, visible, without a single crash in it.
This is the practical answer to "how many hours should I study for CAT". Hours are working. You cannot see them working at a resolution of one mock, and you will drive yourself into the ground trying. Five mocks is roughly the window at which my own effort became visible above the noise.
Question four: where did the hours actually go, and where did they land?
Overall scores hide the most important thing in the log.
| Section | Trend / 30 days | r-squared | First 5 | Last 5 | SD |
|---|---|---|---|---|---|
| VARC | +3.85 | 0.124 | 16.6 | 24.8 | 10.8 |
| QUANT | +2.82 | 0.130 | 26.8 | 35.6 | 7.7 |
| DILR | +0.16 | 0.000 | 18.4 | 20.6 | 7.5 |
Three months of consistent logged hours and DILR has an r-squared of 0.000 against time. Not a weak trend. No trend. Two marks of movement between the first five mocks and the last five, and a fitted slope of a sixth of a mark per month.
This is the single most useful thing in the whole dataset, and it reframes the hours question entirely. The binding constraint on my DILR was never how many hours I put in. It was that the hours were going into solving sets, when the thing costing me marks was choosing which set to attempt - a different skill that I was not practising at all. On the 30 July mock where I scored 2 in DILR, my own note from that week reads that 30 marks were available on the paper. The knowledge was there. The selection was not.
No amount of additional study time fixes a problem you have misdiagnosed. That is why "how many hours" is the wrong first question and "which section has stopped moving" is the right one.
So what is the actual answer on hours?
If you want a number from my log, here is one, with every caveat attached: at a rate of 21 hours a week, the fitted trend of 0.228 marks per day works out to roughly 7.6 marks per 100 hours studied, on a paper scored out of 198.
Do not carry that number around. It is a slope fitted to 20 points with an r-squared of 0.137, over three months, for one person, in a window where I also changed what I studied several times. It is not a conversion rate and it will not hold for you. Its only job is to make a point that a vaguer sentence cannot: the return on an hour is small, slow, and completely invisible against day-to-day variance. Anyone quoting you a confident hours-to-percentile mapping has not checked it against a log.
The useful version of the answer:
- Enough hours to sit a mock a week and analyse it properly. In my log, the 7-day gaps are the ones that left room for a real analysis pass; the 1-2 day gaps taught me almost nothing because there was no time between them to act on anything.
- Enough consistency that a five-mock window means something. Five mocks is where my own signal cleared the noise. Fewer than that and you are reading weather.
- Distributed by section deficit, not by comfort. QUANT rewarded me most reliably, so QUANT is where I kept spending, and my weakest section stood still for a quarter as a direct result.
Where this is weak
n = 1, and n = 20 mocks. Every figure here describes one candidate over 90 days. With a residual standard deviation of 16.9 marks, twenty observations cannot detect a modest effect even if one is there - so "no relationship found" here is genuinely different from "no relationship exists".
Elapsed days is a proxy for hours, not a measurement of them. I do not have a clean per-week hour total lined up against each mock, so the 0.137 is a fit of score against the calendar. A properly specified version - actual hours in the seven days before each mock, against that mock's score - would be a better test and I cannot run it from published data.
The mocks are not a controlled instrument. They come from two providers across three series plus three actual CAT 2020 papers, with different difficulty and different cohorts. Some of the 16.9 marks of residual scatter is paper difficulty rather than me. That cuts both ways: it inflates the noise, and it is also the reality any aspirant is reading their own scores through.
What I studied changed during the window. The trend fits a straight line to a period in which I was actively reallocating effort. A rising QUANT and a flat DILR is partly a story about where the hours went, not just about how many there were.
I built the tracker this data came from, so treat the enthusiasm accordingly. I have tried to publish the findings that argue against tracking as loudly as the ones that support it, which is why the flat DILR line and the r = 0.036 are the headline rather than a footnote.
The one sentence to keep
My overall score nearly doubled over three months while my weakest section did not move at all, and the number of hours I put in predicts neither of those facts well enough to plan on. Count the sections, not the hours - and read them five mocks at a time.
Karma Yogi stores a mock with its per-section scores and every analysis pass attached to it, so a section that has stopped moving shows up as a flat line instead of as something you have to remember. Free, and the same structure works for JEE and NEET sections.
End of essay
- Anish Guruvelli