← The Casebook

Channel · confidence vs. evidence

Illusion of Validity

The illusion of validity is the unjustified confidence people feel in a prediction when the available evidence fits a coherent, representative pattern, even when that evidence has little or no genuine predictive power.

claimed certaintyoriginthe claimKahneman et al. 1973Oskamp 196533%Bushyhead et al. 1981mechanismsettled · mixed▲ peak
Read it left to right like a beam. The bias’s claim spikes above the claimed-certainty line; each study is a labelled measurement, and the evidence sweep pulls the trace back down. Where it finally settles — and whether it still jitters — is the replication status, not the hype.

Free-run conditions

Kahneman and Tversky (1973) coined the term to describe how confidence in a judgment tracks how well the evidence "fits" a stereotype or story, rather than how diagnostic the evidence actually is. So a scanty, redundant, or unreliable description that paints a vivid picture produces high confidence, and that confidence survives even when the predictor knows statistically that the judgment is near-worthless. Kahneman's canonical illustrations are his own Israeli-army officer-selection work (confident leadership predictions that barely correlated with later outcomes) and an analysis of 25 wealth advisers over 8 years whose year-to-year performance rankings correlated on average about .01.

Channel readings

Kahneman & Tversky (1973) — prediction & confidence studies. Confidence in predictions depended primarily on the representativeness (fit) of the evidence to the predicted outcome, 'with little or no regard for the factors that limit predictive accuracy' — the named illusion of validity.

Daniel Kahneman, Amos Tversky, 1973 · Multiple groups of student/respondent participants (sample sizes vary by sub-study; not summarized as one N)

33%

Oskamp (1965) — overconfidence in case-study judgments. As more information was supplied, judges' confidence rose substantially while accuracy stayed roughly flat near chance — the classic dissociation of confidence from accuracy that underlies the illusion of validity.

Stuart Oskamp, 1965 · N = 32 (8 psychologists, 18 grad students, 6 undergraduates)

Bushyhead & Christensen-Szalanski (1981) — feedback and the illusion of validity in a clinic. Physicians relied on clinical attributes tied to the cases for which they happened to order X-rays rather than attributes predictive of all pneumonia cases; biased, incomplete feedback sustained a persistent but poorly-calibrated confidence — attributed to the illusion of validity.

James B. Bushyhead, Jay J. J. Christensen-Szalanski, 1981 · Physicians at one medical outpatient clinic (field setting)

Why the signal misleads

In Kahneman and Tversky's account the illusion of validity is a by-product of the representativeness heuristic: people predict the outcome that best matches (is most representative of) the input, and read the felt sense of a good match as a signal of accuracy. Because the subjective feeling of confidence is generated by the coherence and fit of the story rather than by its diagnostic value, confidence is largely blind to base rates, the reliability of the data, and regression to the mean. Redundant or correlated inputs amplify the effect: many consistent-but-correlated cues feel like converging evidence and inflate confidence, even though they add little independent information. Kahneman later folded this into System 1 / WYSIATI ('what you see is all there is'): the mind builds the best story from available evidence and ignores what is missing.

It is debated whether 'overconfidence' phenomena reflect a genuine cognitive bias or are partly artifacts of task construction. Gigerenzer and colleagues argued much laboratory overconfidence (and the 'hard-easy effect') stems from unrepresentative item selection and disappears under ecologically sampled questions and frequency formats; others (e.g., Griffin & Tversky) treat the strength-vs-weight of evidence as the core driver. So the empirical pattern is robust but its interpretation as a unitary irrational bias is contested.

Calibration verdict

Partial lock — robust in its narrow form, unresolved in its broad one; the trace still jitters.

The core observation — that subjective confidence is driven by evidence fit/coherence and can be sharply decoupled from accuracy — is well replicated (Oskamp 1965; the broader confidence-accuracy and overconfidence literatures; the robust superiority of actuarial over clinical judgment per Dawes, Faust & Meehl 1989 and the Grove & Meehl meta-analyses). However, 'illusion of validity' as a named, separately-measured effect has little dedicated replication; most evidence comes from adjacent overconfidence/representativeness research. The interpretation of overconfidence as a unitary bias is genuinely contested (Gigerenzer's ecological-sampling and hard-easy-effect critiques), so the underlying mechanism — not the phenomenon — is where the dispute lies.

Recorded over-runs

  • Israeli army officer-selection 'leaderless group' assessments · c.1955 (recounted 2011)

    As a young psychologist in the Israeli Defence Forces (early 1950s), Kahneman and colleagues watched recruits in a leaderless obstacle-field task and made confident predictions about who would make good officers. Statistical follow-up showed the predictions correlated only weakly with actual officer-school performance, yet the assessors' confidence in each fresh judgment was undented — Kahneman's origin story for the term.

  • Wealth advisers' year-to-year performance rankings (~.01 correlation) · reported 2011

    Invited to advise a wealth-management firm, Kahneman analyzed a spreadsheet of 25 advisers' investment results over 8 consecutive years, computing 28 between-year correlations of their performance rankings. The average correlation was about .01 — consistent with luck, not skill — yet the firm paid bonuses for skill and the advisers remained confident. The firm carried on unchanged after he presented the result.

  • Physicians ordering chest radiographs for suspected pneumonia · 1981

    In a real outpatient clinic, physicians ordered chest X-rays selectively, learning chiefly from the cases they chose to image. This biased, incomplete feedback left them confidently relying on clinical cues that did not predict pneumonia across all cases — a documented field instance of feedback sustaining the illusion of validity.

Damping

Anchor confidence to documented predictive accuracy, not narrative fit: check base rates and track records, prefer simple statistical/actuarial models over holistic intuition, ask whether your 'multiple cues' are actually independent or merely redundant, and seek the missing or disconfirming evidence (counter WYSIATI). Calibration training and outcome feedback on full samples (not just the cases you acted on) help.

Actuarial models reliably beat confident clinical prediction across domains (Dawes, Faust & Meehl 1989; Grove & Meehl 1996 meta-analysis), and biased partial feedback is what sustained physicians' miscalibrated confidence (Bushyhead & Christensen-Szalanski 1981).

Reading the trace in the wild

You feel very sure of a prediction mainly because the evidence 'hangs together' as a neat story or strongly matches a type — and that confidence doesn't drop when you remind yourself the data are thin, redundant, or known to predict poorly. Tell-tale signs: confidence rising as you accumulate correlated details, ignoring base rates and track records, and conviction that survives disconfirming statistics.

First measured by Daniel Kahneman, Amos Tversky, 1973 — On the Psychology of Prediction.

Adjacent channels

File your own case

Open the same case on your own draft.

Paste a memo, a research draft, or a strategy argument. It is scored against all 175 cards, and the strongest two or three risks come back with the evidence quoted and one practical next check.

Open a case on your draft →