← The Casebook

Calibration Certificate No. hot-hand-fallacy

Hot-Hand Fallacy

The hot-hand fallacy is the belief that a streak of successes raises the probability of further success — that a person or process is "hot" — beyond what the evidence supports.

calibratedintuitioncalibrated
systematic

documented offset from the normative answer

Status — contestedgrounding ± strong

The original GVT 'no hot hand' result was treated as robust for 30 years, but Miller and Sanjurjo (2018, Econometrica) proved a streak-selection bias in the standard estimator that, once corrected, reverses GVT's own data into a sizeable positive effect (~+13 pp). Independent studies controlling for shot difficulty (Bocskocsky, Ezekowitz & Stein 2014: 1.2–2.4 pp) and in volleyball (Raab, Gula & Gigerenzer 2012) also find small real hot hands. The behavioral claim that people sometimes over-believe in streaks survives, but the strong claim that the hot hand is a pure illusion does not. Earlier meta-analytic support (Czienskowski et al.) inherits the same bias and is therefore discounted.

Calibration record

GVT controlled shooting + betting studyRaw average difference about +3 to +4 percentage points (reported as non-significant); 76ers serial correlations mostly near zero or slightly negativeNo significant positive serial correlation; the average within-player difference between P(hit|prior hits) and P(hit|prior misses) was near zero (about +3 to +4 percentage points raw, not significant by their test). Bets did not predict outcomes. Authors concluded the hot hand is illusory. Gilovich, Vallone & Tversky, 1985 · ~26 Cornell players (100 shots each); 9 76ers players; Celtics free-throw records
Streak-selection bias proof + reanalysis of GVTBias ≈ -8 pp (n=100, k=3); corrected GVT hot-hand effect ≈ +13 pp (p<.01); +10 pp at k=4For a 100-shot sequence and streak length k=3 the bias is about -8 percentage points. Correcting GVT's data flips the conclusion: the average bias-adjusted hot-hand difference rises from about +3 to roughly +13 percentage points (p<.01); for k=4 about +10 points; a reanalysis of GVT's betting data shows shooters hit about +7 points more when bettors predicted a hit. Joshua B. Miller & Adam Sanjurjo, 2018 · Same GVT Cornell dataset, reanalyzed
Hot hand with optical tracking (shot-difficulty controls)≈ 1.2 to 2.4 percentage points increased make probabilityPlayers exceeding recent expectations take harder, more contested shots yet still convert slightly above expectation — a small but real hot-hand effect. Andrew Bocskocsky, John Ezekowitz & Carolyn Stein, 2014 · ~83,000 shots, 2012–13 NBA season
Hot hand in volleyball and allocationnot reported as a single pooled estimate (significant for a subset of players)A hot hand was present for a subset of players; playmakers detected and exploited it adaptively, increasing team hits. Markus Raab, Bartosz Gula & Gerd Gigerenzer, 2012 · Professional volleyball players (German Bundesliga data)

Source of systematic error

GVT attributed the supposed fallacy to misperception of randomness: people expect short random sequences to be more alternating than they really are (a consequence of Tversky and Kahneman's 'belief in the law of small numbers' and the representativeness heuristic), so genuinely chance streaks look meaningful and are read as evidence of being 'hot'. Memory bias reinforces this — vivid streaks are recalled more readily. The Miller–Sanjurjo correction, however, shows the deeper issue is statistical, not purely psychological: the standard streak-conditioned estimator is itself biased toward finding no hot hand, so part of what GVT diagnosed as a human error was an error in their own measure.

(1) Pure-illusion account (GVT 1985): belief in continuation is a cognitive bias with no real signal. (2) Statistical-artifact account (Miller & Sanjurjo 2018): the apparent absence of a hot hand was largely a selection-bias artifact; modest hot hands are real, so the 'fallacy' may be overstated. (3) Ecological/adaptive account (Raab et al. 2012): detecting streaks can be a rational, useful cue in settings where the hot hand genuinely exists. The mechanism debate is unresolved and partly definitional — whether 'the fallacy' means believing in any continuation vs. over-estimating its size.

Recalibration procedure

State the base rate and sample size first, then ask what a purely random process of that length would produce before attributing a streak to 'heat.' For any decision that conditions on a streak (bet more, allocate to the hot fund, keep feeding the hot shooter), check for streak-selection bias — the naive 'success rate after a streak' is downward-biased, so neither a high nor a low number should be taken at face value without the Miller–Sanjurjo correction or a permutation test.

Miller & Sanjurjo (2018) show the standard streak-conditioned estimator is biased by roughly -8 percentage points for a 100-trial sequence (k=3), large enough to flip GVT's headline conclusion; correcting it recovers a ~+13-point effect, demonstrating that the cure for the fallacy is explicit statistical correction, not intuition.

Cross-calibrated against

Calibrated by Thomas Gilovich, Robert Vallone, Amos Tversky, 1985 — The hot hand in basketball: On the misperception of random sequences. Cognitive Psychology, 17(3), 295–314.

Uncertainty: Effect sizes for the Miller–Sanjurjo reanalysis (-8 pp bias; +13 pp corrected GVT effect at k=3; +10 pp at k=4; +7 pp on the betting data) were read directly from a hosted full-text of the Econometrica paper and are high-confidence; the original GVT 'raw' difference (~+3 to +4 pp) is likewise from that text. The Bocskocsky–Ezekowitz–Stein 1.2–2.4 pp figure comes from a search summary of their MIT Sloan Sports Analytics Conference paper; I could not fetch the exact PDF, so treat that range as well-sourced but secondary, and the URL points to the conference paper repository rather than a stable DOI. The Raab et al. volleyball result is confirmed via PubMed but I did not extract a single pooled effect size. Two of the three documented real-world cases (Chen-Moskowitz-Shue; Croson-Sundali) primarily document the gambler's fallacy, the sibling bias — included because they share the same law-of-small-numbers mechanism and are the best-documented field evidence on streak (mis)perception, but they are not pure hot-hand demonstrations. The overall framing (contested, not robust) is well supported; the precise boundary between 'a real small hot hand' and 'over-belief in it' remains genuinely unsettled in the literature.

File your own case

Open the same case on your own draft.

Paste a memo, a research draft, or a strategy argument. It is scored against all 175 cards, and the strongest two or three risks come back with the evidence quoted and one practical next check.

Open a case on your draft →