Karma YogiKarma Yogi
Start
The Journal
Study Techniques

10 September 2026
12 min read

A
Anish Guruvelli
Karma Yogi

Best Time to Study Morning or Night: What 3.4M Logins Say

There is no universal best hour, and the largest datasets on this question disagree about which direction the effect even runs. Here is what four studies actually found, including the one whose own prediction was refuted, and a three-week test you can run on your own log instead.

There is no universal best hour, and the two largest datasets on this question point in opposite directions. An analysis of 3.4 million learning-management logins from 14,894 university students found that late chronotypes scored worse at every class time, and that grades improved later in the day for everyone - the opposite of what "study at your peak" predicts. A separate analysis of 2,034,964 Danish test results found scores falling 0.9% of a standard deviation for every hour later the test was taken, while a 20-to-30-minute break recovered 1.7% - close to two hours of decline, bought back with half an hour off. My own log says something duller and more useful: across 340 sessions my mean is 59 minutes, and the sessions I rate worst are the late-evening ones squeezed in after work.

So the honest answer is: pick the hour you can defend against the rest of your life, take breaks inside it, and stop reading articles about this one. What follows is the actual state of the evidence, including the parts that contradict each other, and a three-week test that will tell you more about you than any of it.

Why "are you a lark or an owl" is the wrong first question

The pop version of this topic runs: chronotypes are fixed, find yours, study at your peak. The research idea underneath it is real and is called the synchrony effect - the claim that people perform better on cognitive tasks at the time of day that matches their circadian preference. It has genuine experimental support. It also failed the largest real-world test anyone has run on it, and that failure is more interesting than the original finding.

The four studies below are the strongest ones I could open and read in full. They use four different populations, four different outcome measures, and they do not agree. That disagreement is the finding.

The 3.4-million-login study, and the prediction it refuted

Smarr and Schirmer, Scientific Reports volume 8, published March 2018. They took 3.4 million login events from the learning management system at Northeastern Illinois University across four semesters (Fall 2014 to Spring 2016), covering 14,894 students, and inferred each student's chronotype from the median phase of their activity on non-class days.

That produced three groups: larks (3,857 students, median activity time 12:46), finches (11,261, 16:22) and owls (3,419, 20:20). Sixty per cent of students showed average daily social jet lag - a gap between their class-day and free-day rhythms - of more than 30 minutes; only 40.4% stayed inside 30 minutes.

Their stated hypothesis was the synchrony prediction: owls should improve across the day, larks should worsen. It was refuted. What they found instead:

  • Everyone did better later. Average evening class grade points ran at least 0.27 points higher than average morning ones, for all three chronotypes.
  • Owls were worse at every hour. After normalising grades by class time to strip out the time-of-day effect, owls still showed a significant deficit at all class times (chi-squared = 114, p = 2 x 10 to the -25).
  • And owls took the most morning classes (chi-squared = 13.82, p = 0.001), which is a scheduling problem stacked on top of a performance one.

The authors are careful about what this is. LMS logins are "not directly measuring biological outputs like sleep or circadian melatonin level" - they measure time spent on academically targeted effort. Login frequency was too low to build a daily actogram per student. And they flag that more than half the classes at that university met only twice a week at variable start times, which may amplify the owl disadvantage relative to a school with a fixed timetable.

The 2-million-test study says the clock beats the chronotype

Sievertsen, Gino and Piovesan, PNAS 113(10):2621-2624, February 2016. This is the biggest dataset in this post: 2,034,964 test observations from 570,376 Danish students aged 8 to 15, across school years 2009/10 to 2012/13, on a compulsory national computer-based test that schools scheduled at whatever hour suited them.

Two numbers matter, and the second one is the useful one:

  • For every hour later in the day, test scores fell by 0.9% of a standard deviation (95% CI 0.7-1.0%). The authors put that at roughly the same size as 1,000 USD lower household income, a month less parental education, or 10 school days.
  • A 20-to-30-minute break before the test improved performance by 1.7% of a standard deviation (95% CI 1.2-2.2%) - about 19 school days, or roughly double the damage that two hours of elapsed day does.

The decline was not linear. Scores deteriorated from 8:00 to 9:00, improved from 9:00 to 10:00, and kept alternating through the day in a pattern that tracked when breaks fell. That is the shape of fatigue and recovery, not the shape of a circadian curve.

The transfer caveat is large and the authors state it themselves: "the students in our sample are young children and early adolescents, and older adolescents may fare differently." An eight-year-old sitting a school test is not a working adult sitting a 40-minute CAT section, and nobody has run this study on the second group. What I take from it is not the 0.9% number but the ratio: in the only dataset here with two million observations, the break mattered more than the hour.

Where the synchrony effect does hold

Goldstein, Hahn, Hasher, Wiprzycka and Zelazo, Personality and Individual Differences 42(3):431-440, 2007. This is the study most often cited as proof that you should study at your peak, so it is worth reading what it actually tested.

Eighty adolescents aged 11-14 (mean 12.48), screened down from 259 telephone interviews and classified by the Children's Morningness-Eveningness Preferences scale into the outer quartiles. A 2 x 2 design - Morning or Evening type, tested 8-10am or 1-3pm - with 20 per cell.

The result splits in a way that matters more than the headline:

  • Fluid intelligence showed the effect. Tested at their optimal time, adolescents scored 12.08 (SD 2.49) against 10.86 (SD 2.24) at their non-optimal time: t(78) = 2.29, p = .03, d = 0.52, which the authors describe as about a six-point difference in Full Scale IQ equivalents.
  • Crystallised intelligence showed nothing. On Vocabulary, F was less than 1, not significant. No synchrony effect at all.

That split is the single most useful thing in this post for a CAT aspirant, and it is also where I have to be honest about how far it stretches. The effect appeared on reasoning and working memory, and was absent on stored vocabulary knowledge. If it transfers to CAT at all, it transfers to the parts of the paper that are live reasoning under time pressure - DILR sets, unfamiliar Quant - and not to the parts that are retrieval of things you already know. That is my inference from the measures used, not a finding. Nobody has tested a synchrony effect on a CAT-shaped paper.

The authors' own limits: the study ran in summer, when adolescents have more control over their sleep than in term time, so they argue the effect may be underestimated here; and sleep duration was self-reported rather than measured.

The closest thing to a randomised test

Rodriguez Ferrante, Goldin, Sigman and Leone, npj Science of Learning volume 8, June 2023. Small sample, unusually strong design: 407 students in a Buenos Aires secondary school where the school shift is assigned by lottery at the start of secondary school - morning, afternoon or evening. Chronotype measured with the Munich Chronotype Questionnaire, using sleep-corrected mid-sleep on free days.

Odds of grade retention (being held back), per one hour later chronotype:

School shiftOdds ratio per hour later chronotypeReading
Morning1.65666% higher odds of being held back
Afternoon1.168Not statistically significant
Evening0.949Slightly protective

So a late chronotype cost these students something in the morning shift and nothing in the later ones. That is the synchrony account surviving where the Chicago dataset killed it - on a much smaller sample, with a much better allocation mechanism. The authors' caveats: grades were assigned by teachers who were not blind to the students, chronotype was self-reported, and they say plainly that the results "are based on correlations, which do not allow us to establish causality relationships" despite the lottery.

The four studies side by side

Study Sample What it measured What it found
Goldstein et al. 2007 80 adolescents Lab reasoning and vocabulary, 11-14 year-olds Synchrony effect on fluid reasoning (d = 0.52); none on vocabulary
Sievertsen et al. 2016 2,034,964 tests National school tests, ages 8-15, Denmark -0.9% SD per hour later; +1.7% SD from a 20-30 min break
Smarr & Schirmer 2018 14,894 students Semester grades vs class time, US university Everyone better later; owls worse at every hour; synchrony prediction refuted
Rodriguez Ferrante et al. 2023 407 students Grade retention by lottery-assigned school shift Late types penalised in the morning shift only (OR 1.656)

What the disagreement actually means

There are three competing accounts here, and each of these datasets supports a different one:

  1. Synchrony. You perform best at your preferred time. Supported by Goldstein (lab, reasoning tasks, 80 people) and by the morning-shift half of the Buenos Aires result.
  2. Fatigue. Performance declines through the day for everybody, and rest resets it. Supported by two million Danish tests, where the break effect was nearly twice the size of two hours of elapsed day.
  3. Trait. Being a late chronotype is a marker for something else - accumulated sleep debt, an unstable schedule, less protected study time - that costs you at every hour, not just the wrong ones. Supported by the Chicago dataset, whose owls were worse at 9am and at 7pm alike.

You cannot settle this by picking the biggest sample, because the three studies measure different outcomes: a lab reasoning score, a single school test, and a semester of grades. Anyone who tells you the science says to study at 5am, or at 11pm, is quoting one of these three and quietly not mentioning the other two.

What my own log says, and what it cannot say

From my own tracked CAT prep: 340 sessions, 20,100 minutes, 135 active days, a mean session of 59 minutes, a longest session of 190 minutes (a weekend full mock), and a mean self-rated mood of 3.60 out of 5. The weekday shape is a 45-to-60-minute block before work and a 30-to-45-minute review in the evening; the weekends carry the full mocks. The evening sessions are consistently the ones that drag the mood average down, and the protected weekend blocks are the ones that lift it. I wrote about the fragmentation problem behind that shape in the full-time-work post.

Here is what my log cannot tell me, and it is the same hole in every "I'm a morning person" claim you will read: I do not know whether my mornings are better because mornings are better, or because my mornings are protected and my evenings are whatever is left after a working day. There is no control condition in a personal log. The variable I actually changed was not the hour - it was how much of my attention the rest of the day had already spent.

The three-week test that beats all of this

Rather than picking a camp, run the comparison on yourself. This is deliberately narrow, because a wide version of it produces noise you cannot read:

  1. Pick one activity, not "studying". A 30-question timed Quant set, or one DILR set of the same difficulty band. It has to produce a score.
  2. Pick two slots you can genuinely hold - say 6:30am and 9:30pm - and alternate them week by week, not day by day. Alternating daily confounds the slot with the previous night's sleep.
  3. Log the score, not the feeling. Mood tracks how hard something felt, and difficulty and learning come apart routinely. Score is the thing you are optimising.
  4. Run three weeks minimum and compare the medians, not the best day in each slot.
  5. Then stop. If the gap is smaller than the spread within each slot, the hour is not your constraint and you should go and fix something that is.

Three weeks of one person's data is a weak experiment. It is also a better guide to your own schedule than a Danish primary-school average, because it is measured on the population of one that you are actually optimising.

What I would actually do for CAT

  • Put your hardest section in your most protected slot, not your theoretically peak one. Protection is measurable; peak is not.
  • Take a real break every 40-50 minutes. This is the one recommendation with a two-million-observation dataset behind it, and its effect there was larger than the whole time-of-day effect. My mean session is 59 minutes, which is close to that boundary by accident rather than design.
  • Sit full mocks in the real slot. CAT runs in fixed slots, so the useful adaptation is not "study when you feel sharpest" but "be sharp when the paper is". Move your mocks to the slot you expect, months before, and let the trend tell you if it costs anything.
  • Fix sleep before you fix the hour. The chronotype literature and the sleep literature keep converging on the same point, and the sleep effect is much better established - the sleep post has that evidence.
  • Stop optimising once you have picked. The cost of switching your schedule every fortnight is certain; the benefit of finding the perfect hour is contested across four studies.

Where this is weak

  • Not one of these studies tested an adult sitting a CAT-shaped paper. The populations are 11-14 year-olds in a lab, 8-15 year-olds sitting national school tests, US undergraduates, and Argentine secondary students. Every number here is a description of a different population, and I am extrapolating in all four cases.
  • The direction of the effect is genuinely contested, not merely uncertain in size. One dataset says later is better for everybody; another says later is worse for everybody. Both are large. I cannot reconcile them and I have not seen anyone who can.
  • Chronotype was self-reported in three of the four. Only the Chicago study used behavioural data, and it inferred chronotype from web logins, which the authors say is a measure of academic effort rather than biology.
  • The synchrony study is small. Eighty participants, 20 per cell, one lab session each. A d of 0.52 from that design is a signal worth following, not a settled quantity.
  • My own log has no control. It shows what my schedule was, not what it would have produced if reversed, and I have never run the alternate-week test I am recommending to you.
  • Grade retention is not a mock score. The Buenos Aires outcome is being held back a year, which is dominated by factors far larger than the hour a class starts.

Sources

  • Goldstein, D., Hahn, C. S., Hasher, L., Wiprzycka, U. J., & Zelazo, P. D. (2007). Time of day, intellectual performance, and behavioral problems in Morning versus Evening type adolescents. Personality and Individual Differences, 42(3), 431-440. Read in full via PMC.
  • Sievertsen, H. H., Gino, F., & Piovesan, M. (2016). Cognitive fatigue influences students' performance on standardized tests. PNAS, 113(10), 2621-2624. Read in full via PMC.
  • Smarr, B. L., & Schirmer, A. E. (2018). 3.4 million real-world learning management system logins reveal the majority of students experience social jet lag correlated with decreased performance. Scientific Reports, 8, 4793. Read in full via PMC.
  • Rodriguez Ferrante, G., Goldin, A. P., Sigman, M., & Leone, M. J. (2023). A better alignment between chronotype and school timing is associated with lower grade retention in adolescents. npj Science of Learning, 8, 21. Read in full via PMC.
  • Session log: 340 sessions, 20,100 minutes, 135 active days, 59-minute mean session - first-party, from my own Karma Yogi account.

If you want to answer this for yourself rather than for the average Danish nine-year-old, the thing you need is a per-session record with a timestamp and a score attached. Karma Yogi logs both, and the heatmap shows which hours you actually study rather than which ones you believe you do.

End of essay

- Anish Guruvelli

Common questions

Is morning or night better for studying?
Neither, universally. The largest datasets disagree in direction: 3.4 million university logins showed grades improving later in the day for every chronotype, while two million Danish school tests showed scores falling 0.9% of a standard deviation per hour later. Pick the slot you can protect and hold it.
What is the synchrony effect?
The claim that you perform better on cognitive tasks at the time of day matching your circadian preference. It has real support - a 2007 study of 80 adolescents found a d of 0.52 on fluid reasoning - but the largest real-world test of it, on 14,894 students, refuted its own prediction.
Do night owls actually do worse academically?
In one large dataset, yes, and not only in morning classes. Smarr and Schirmer found late chronotypes underperformed at every class time even after normalising for time of day. That points at something travelling with late chronotype - sleep debt, schedule instability - rather than at a simple mismatch.
How many hours after waking is your brain sharpest?
There is no reliable number for this, and the studies I could read do not support one. The Danish data shows an alternating pattern of decline and recovery through the day that tracks break timing, not a single peak you can schedule around.
Should I study CAT quant in the morning or evening?
Put it wherever you can protect a genuinely uninterrupted block. The one hint from the literature is that the synchrony effect appeared on fluid reasoning and not on stored vocabulary, so if any part of CAT is time-of-day sensitive it is live reasoning - but nobody has tested this on a CAT paper.
Do breaks matter more than what time I study?
In the only two-million-observation dataset here, yes. A 20-to-30-minute break was worth 1.7% of a standard deviation, against 0.9% lost per hour later in the day. That break effect is the most transferable finding in this whole literature, and the cheapest to act on.
Can I change my chronotype to become a morning person?
None of the four studies here tested that, so I cannot answer it from evidence I have read. What they do show is that the cost of being a late type appears mainly when your schedule forces early performance, which is an argument for changing the schedule where you can.
Does the time I take mocks matter?
More than the time you study, in my view. CAT runs in fixed slots, so the useful adaptation is being sharp when the paper is. Move your full mocks to the slot you expect months ahead and watch whether the trend costs you anything.
How do I find my own best study time?
Alternate two fixed slots by week - not by day, which confounds the slot with last night sleep - running the same scored activity in each, for at least three weeks, then compare medians. If the gap is smaller than the spread inside each slot, the hour is not your constraint.
Why does my own log not settle this?
Because a personal log has no control condition. Mine shows mornings going better than evenings, but my mornings are protected and my evenings are whatever survives a working day, so the variable that differs is not the hour. That is the same flaw in every anecdotal claim about this.