← The Casebook

Calibration Certificate No. insensitivity-to-sample-size

Insensitivity to Sample Size

Insensitivity to sample size is the tendency to judge the probability of obtaining a sample statistic without adequately considering the sample size, treating small samples as equally representative of populations as large samples.

calibratedintuitioncalibrated
systematic

documented offset from the normative answer

Status — robustgrounding ± strong

The core finding—that people fail to adequately weight sample size in probability judgments—has been replicated extensively across populations including trained psychologists and statisticians. Tversky and Kahneman's hospital problem has been replicated with similar results (approximately 20% correct). Zhan et al. (2022) found people do not differentiate findings from samples varying by a factor of 100 in between-participant designs. The bias persists even among individuals with statistical training, suggesting it is a deeply ingrained cognitive tendency rather than a knowledge deficit.

Calibration record

Hospital births problemnot reportedOnly 22% of subjects correctly identified the smaller hospital. Fifty-six percent said the number of such days would be 'about the same,' and 22% chose the larger hospital, demonstrating failure to appreciate that sampling variance decreases with sample size. Tversky & Kahneman, 1974 · n = undergraduates (exact N not reported in available sources); 22% correct response rate
Mean height sample judgmentnot reportedSubjects assigned approximately the same probability to obtaining the extreme mean height regardless of whether the sample contained 10, 100, or 1,000 men, demonstrating that they failed to incorporate sample size into their probability judgments. Tversky & Kahneman, 1971 · n not reported
Replication with training interventionnot reportedFound evidence that people do believe in the empirical law of large numbers when large and small samples are juxtaposed directly, but within-participant designs may suffer from experimenter demand effects. Pure between-participant designs free from demand effects show people tend to operate on the law of small numbers, not differentiating findings from samples varying by a factor of 100. Sedlmeier & Gigerenzer, 1997 · Multiple experiments; exact N varies

Source of systematic error

Insensitivity to sample size arises from the representativeness heuristic: people intuitively judge samples by how similar they appear to the parent population, without considering the statistical principle that larger samples more closely approximate population parameters. People focus on the content or story of the data (e.g., '60% boys') rather than the reliability information embedded in sample size. This reflects a deeper tendency to construct coherent narratives from limited evidence while neglecting uncertainty indicators.

Some researchers argue that when small and large samples are directly juxtaposed, people do show sensitivity to sample size (the 'empirical law of large numbers'; Sedlmeier & Gigerenzer, 1997). However, between-participant designs free from demand effects suggest this sensitivity is often an artifact of experimental context.

Recalibration procedure

Always ask 'How large is the sample?' before accepting statistical claims. Use confidence intervals rather than point estimates to appreciate uncertainty. For important decisions, require larger samples or meta-analytic evidence aggregating across multiple studies.

Training using simulation-based methods (active experience sampling from populations of different sizes) shows modest improvements in sensitivity to sample size, though the bias is resistant to purely didactic instruction.

Cross-calibrated against

Calibrated by Amos Tversky, Daniel Kahneman, 1971 — Belief in the law of small numbers. Psychological Bulletin, 76(2), 105–110.

Uncertainty: Some debate exists about whether sample size neglect truly reflects cognitive limitations or is partly an artifact of experimental design (within-participant vs. between-participant). The exact percentages from the hospital problem vary slightly across replications.

File your own case

Open the same case on your own draft.

Paste a memo, a research draft, or a strategy argument. It is scored against all 175 cards, and the strongest two or three risks come back with the evidence quoted and one practical next check.

Open a case on your draft →