The "66 days to build a habit" number is a median from 39 people, and the range around it ran from 18 to 254 days. It comes from Lally, van Jaarsveld, Potts and Wardle's 2010 study in the European Journal of Social Psychology, which followed 96 volunteers for 84 days as they repeated one self-chosen eating, drinking or exercise behaviour daily. Of the 82 who supplied enough data, an asymptotic habit curve was a good fit for 39 - 48%. The other half never became consistent enough to form a habit at all, inside a paid study they had volunteered for.
None of them were studying. That matters more than the number does, and the rest of this post is about what genuinely transfers to exam prep and what does not - including the finding that makes a streak counter defensible and the one that makes most streak designs wrong.
What the study actually did
101 university students in London attended an initial meeting; 96 enrolled, 30 men and 66 women, almost all postgraduates. Each chose one healthy behaviour to perform daily in a consistent context - "after breakfast", "before dinner" - for 84 days, and logged onto a website each day to complete the Self-Report Habit Index automaticity subscale. They were paid £30 for participating, not contingent on succeeding.
Fourteen stopped entering data before day 60 and were treated as dropouts. The remaining 82 logged on a median of 47 of the 84 days. A nonlinear curve was then fitted to each person's automaticity scores. It converged for 62 people; after excluding poor fits and unrealistic asymptotes, the model was a good fit for 39 of 82. Here are the parameters for those 39, straight from the paper's Table 2.
| Curve parameter (N = 39) | Median | Quartiles (Q1:Q3) | Minimum | Maximum |
|---|---|---|---|---|
| Days to reach 95% of asymptote | 66 | 39:102 | 18 | 254 |
| Asymptote (max automaticity, scale max 42) | 35 | 29:43 | 21 | 48 |
| Rate of change (c) | 0.042 | 0.028:0.064 | 0.010 | 0.170 |
| Model fit (R²) | 0.88 | 0.84:0.94 | 0.72 | 0.98 |
Three things fall out of that table that the popular version of this study never mentions.
1. The interquartile range alone is 39 to 102 days. Half the successful habit-formers fell outside a span of two and a half months. A single headline number was never the finding; the spread was.
2. The 254 is an extrapolation. The study ran 84 days. Anyone whose modelled asymptote fell beyond that had a curve projected past the data, which the authors say explicitly - high asymptotes "were often due to participants not reaching a plateau within the 84 days". So the top of the range is a model output, not an observation.
3. Half the sample is missing from the headline. 39 of 82 got a good fit. The authors' own gloss: "even in this study where the participants were motivated to create habits, approximately half did not perform the behaviour consistently enough to achieve habit status."
The behaviour-type result, and why 66 is a floor for studying
Among the 39, the chosen behaviours split into eating, drinking and exercise, and the paper compares them.
| Behaviour type | N | Days to 95% asymptote (median) | Q1:Q3 | Compliance |
|---|---|---|---|---|
| Drinking (e.g. a glass of water) | 15 | 59 | 39:75 | 93% |
| Eating (e.g. fruit with lunch) | 10 | 65 | 35:106 | 80% |
| Exercise | 13 | 91 | 44:118 | 86% |
The difference in days was not statistically significant (Kruskal-Wallis, p = .328) and the study was not powered for subgroup analysis, so treat it as a direction rather than a result. But the authors take it seriously, and their reasoning is the load-bearing part for anyone applying this to exam prep:
"It is notable that the exercise group took one and a half times longer to reach their asymptote than the other two groups. Given that exercising can be considered more complex than eating or drinking, this supports the proposal that complexity of the behaviour impacts the development of automaticity."
Now place a CAT study session on that scale. Drinking a glass of water after breakfast is one action with one cue. A study session requires choosing a subject, choosing material, clearing 45 to 90 minutes, sustaining attention through difficulty, and doing it in a context that changes with your week. It is more complex than exercise by a wide margin. If exercise took 1.5x longer than drinking, the honest extrapolation is that 66 days is a floor for a study habit, not a target - and even that is extrapolation, because nobody in this study was studying.
The missed-day finding, which is the one that should shape a streak
Lally et al. tested something almost nothing else in this literature has: what a single skipped day actually costs. They defined a missed opportunity as a day the behaviour was not performed, immediately preceded by three days it was, and found 140 such misses across 55 participants.
| Comparison | Change in automaticity score | Reading |
|---|---|---|
| Day immediately after a missed day | -0.29 | A very small dip on a 42-point scale |
| Before the miss → after resuming (N = 67) | +0.55 | Net gain; not significant (Wilcoxon) |
| Three consecutive performed days | +0.79 | The unbroken comparison |
So a broken day cost roughly a quarter of a point of progress against an unbroken one, and the timing of the miss did not matter either - there was no correlation between how far into the study an omission fell and its effect (r = 0.099, p = 0.246, N = 140). The authors' conclusion: "a missed opportunity did not materially affect the habit formation process."
They then draw the distinction that actually matters. They contrast their result with Armitage (2005), which they describe as finding that lapses did have consequences - but there a lapse was defined as not attending for a whole week. Their reconciliation, which I am quoting rather than paraphrasing because I have not read the Armitage paper myself and am relying on their account of it:
"These findings can easily sit side-by-side and suggest that missing one opportunity does not preclude habit formation, but missing a week's worth of opportunities reduces the likelihood of future performance and hinders habit acquisition."
What this means for a streak counter, including mine
A streak is a product mechanic, not a scientific instrument. But the evidence above gives a fairly clear specification for what a good one would and would not do, and it is worth writing down because most streak designs get the second column wrong.
| Design decision | What the evidence supports | What it does not support |
|---|---|---|
| Counting consecutive days | Consistency predicted habit formation: compliance correlated with model fit (r = .34, p = .035) | Treating the count as a measure of learning |
| Reaction to one missed day | Near-zero real cost (-0.29 points, recovered on resuming) | Framing a single miss as failure or loss |
| Reaction to a missed week | This is the interval that appears to matter | Ignoring it because the daily miss was harmless |
| A target number of days | Nothing. 66 is a median with an IQR of 39-102 | "Do this for 21/30/66 days and it sticks" |
| Same cue, same context | The whole study design rests on context stability | Assuming a habit transfers across changed routines |
I built Karma Yogi's streak and I break it. My own account shows 340 logged sessions across 135 distinct active days, at a 59-minute average session - several runs stitched together, not one unbroken line. The full breakdown of what actually kept me logging is in the streak post. The reason the counter is there is the first row of that table: consistency was the one performance variable that predicted whether a habit curve fitted at all. The reason it is not a punishment mechanic is the second row.
To be explicit, because this is exactly the kind of claim products overreach on: the streak feature is a design choice informed by this literature, not a guarantee that any number of consecutive days will lock in a study habit. No study in this area has tested that, for studying, in anyone.
What I would actually do
- Set a floor low enough that a bad day still clears it. The consistency finding is about performing the behaviour at all, not about duration. A 20-minute session that keeps the run alive is worth more to habit formation than a heroic three-hour session followed by four blank days.
- Anchor it to a fixed cue, not a fixed time budget. Every participant in this study attached their behaviour to a stable context. "After dinner" survives a chaotic week; "two hours in the evening" does not.
- Treat one missed day as noise and one missed week as a signal. That is the single clearest actionable split in the evidence, and it is the opposite of how most people react to a broken streak.
- Expect months, and expect it to be worse for studying than for the study's behaviours. Median 66 days for simple behaviours, 91 for the most complex one tested. Studying is more complex than all three.
- Do not confuse an active streak with productive work. The counter measures showing up. Whether the session was worth anything is a separate question - see what deliberate practice actually requires.
Where this is weak, and where it is routinely overstated
- The headline number rests on 39 people. 96 enrolled, 82 usable, 62 fitted, 39 good fit. Every quotation of "66 days" is quoting a median of 39 postgraduate students in London.
- Nobody was studying. The behaviours were eating fruit, drinking water and short bouts of exercise. Applying the timeline to a 90-minute Quant session is extrapolation across both complexity and domain, and I am flagging it rather than hiding it.
- The measure is self-report. Automaticity was captured by the Self-Report Habit Index, which the authors note had never previously been used repeatedly like this. They found no correlation between number of questionnaire completions and any curve parameter, which is reassuring, but it remains a rating scale rather than an objective measure.
- 254 days was never observed. It is a curve fitted to 84 days of data and projected outward. Quoting it as "it can take up to 254 days" reports a model, not a person.
- The behaviour-type comparison was not significant. p = .328 on the days-to-asymptote difference, and the study was explicitly not powered for it. The "exercise took 1.5x longer" observation is suggestive, and the authors say so; it is the basis of my "66 is a floor" reasoning and that reasoning inherits the weakness.
- Missed-opportunity analysis had limited data. 140 misses, 67 with usable follow-up scores. The authors themselves say more work is needed on whether the timing of a lapse matters.
- My own streak data is one person. 340 sessions and 135 active days on a single CAT account illustrate a shape; they establish nothing about anyone else, and I have not sat the exam yet.
- Nobody has run the study you want. There is no trial of streak mechanics against exam outcomes in Indian competitive prep. If one exists, I have not found it.
Sources
- Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., & Wardle, J. (2010). How are habits formed: Modelling habit formation in the real world. European Journal of Social Psychology, 40(6), 998-1009. doi:10.1002/ejsp.674; open-access copy read in full at the ISPA repository. Every figure above is from its Results, Tables 1 and 2, and Discussion.
- Armitage, C. J. (2005) is referenced above only as Lally et al. describe it in their Discussion. I did not open that paper, and have not verified its numbers independently.
- Session and active-day counts are first-party, from my own Karma Yogi account; the full breakdown is in the streak post.
Karma Yogi counts active days rather than grading them, which is the closest a counter can get to what this evidence supports: consistency is the variable that predicted habit formation, a single missed day was not a measurable setback, and a missed week is the thing worth noticing.
End of essay
- Anish Guruvelli