← The Casebook

Channel · confidence vs. evidence

Hard-Easy Effect

The hard-easy effect is the observed tendency for people's confidence to be too high relative to accuracy on hard tasks (overconfidence) and too low on easy tasks (underconfidence), as if they fail to fully adjust judged confidence to overall task difficulty.

claimed certaintyoriginthe claimLichtenstein et al. 1977Klayman et al. 1999Juslin et al. 2000Merkle 2009mechanismsettled · contested▲ peak
Read it left to right like a beam. The bias’s claim spikes above the claimed-certainty line; each study is a labelled measurement, and the evidence sweep pulls the trace back down. Where it finally settles — and whether it still jitters — is the replication status, not the hype.

Free-run conditions

In calibration studies using general-knowledge questions, average confidence tracks task difficulty far less than accuracy does, so confidence exceeds accuracy on hard item sets and falls short of it on easy item sets. First documented in the 1970s overconfidence literature (Lichtenstein and Fischhoff, 1977), it was long treated as the most consistent finding in laboratory overconfidence research. However, a major line of work (Gigerenzer et al. 1991; Juslin, Winman and Olsson, 2000; Merkle, 2009) argues the pattern is largely a statistical artifact of unrepresentative item selection, scale-end effects, linear dependency and regression to the mean rather than a genuine cognitive bias, and that it nearly vanishes once those artifacts are controlled.

Channel readings

Lichtenstein & Fischhoff (1977) — original demonstration. Overconfidence was larger on hard questions and reduced or reversed (toward underconfidence) on easy questions, i.e., confidence was insufficiently sensitive to difficulty.

Sarah Lichtenstein, Baruch Fischhoff, 1977 · not reliably retrieved

Klayman, Soll, González-Vallejo & Barlas (1999) — difficulty vs. domain. Little overconfidence with two-choice questions but pronounced overconfidence with subjective confidence intervals; over/underconfidence varied systematically with the domain of questions, not as a clean function of difficulty.

Joshua Klayman, Jack B. Soll, Claudia González-Vallejo, Sema Barlas, 1999 · not reliably retrieved

↓ artifact

Juslin, Winman & Olsson (2000) — artifact critique. Very little support for a cognitive-processing bias; near elimination of the hard-easy effect once scale-end effects and linear dependency are controlled; a residual difference between representative and selected samples not reducible to difficulty.

Peter Juslin, Anders Winman, Henrik Olsson, 2000 · meta-analytic re-analysis across multiple prior datasets (exact N not retrieved)

↓ artifact

Merkle (2009) — mathematical conditions. All types of judges exhibit the hard-easy effect in almost all realistic situations; its presence therefore cannot distinguish between judges or support specific confidence models.

Edgar C. Merkle, 2009 · not applicable

Why the signal misleads

Two broad accounts compete. The cognitive account holds that judges anchor on a moderate baseline confidence and adjust insufficiently for how hard or easy a particular item set is, producing overconfidence on hard sets and underconfidence on easy ones. The artifact account holds that the pattern is mostly mechanical: when items are not a representative sample of a natural reference class (experimenters often pick 'tricky' hard items), cue validities shift; and even with unbiased judges, unsystematic error in confidence plus scale-end ceiling/floor effects and the linear dependency between confidence and the bias measure force a negative correlation between difficulty and over/underconfidence (regression to the mean).

Cognitive-bias view (Lichtenstein & Fischhoff tradition) vs. ecological/representative-sampling view (Gigerenzer, Hoffrage & Kleinbölting 1991) vs. pure statistical-artifact view (Juslin, Winman & Olsson 2000; Merkle 2009). Klayman et al. (1999) is intermediate, finding domain matters more than difficulty and that the effect depends heavily on elicitation method.

Calibration verdict

No lock — the trace stays high and unstable; the evidence has not resolved.

The descriptive pattern (confidence exceeds accuracy more on hard sets) replicates extremely reliably and is often called the most consistent finding in laboratory overconfidence research. But its INTERPRETATION as a cognitive bias is strongly contested: Juslin, Winman & Olsson (2000) report the effect nearly disappears once scale-end effects, linear dependency and regression artifacts are controlled, and Merkle (2009) shows it appears for essentially all judges in realistic conditions, so it cannot diagnose a bias. Klayman et al. (1999) find over/underconfidence tracks domain and elicitation method rather than difficulty per se. The honest summary: a robust statistical regularity, a contested psychological bias.

Recorded over-runs

  • Overconfidence as a cause of diagnostic error in medicine · 2008

    A widely cited review argues physician overconfidence and poor calibration contribute to diagnostic error, with confidence often failing to fall as much as accuracy on harder cases; cited as a real-world manifestation of difficulty-insensitive confidence. (Note: this is a review/argument linking overconfidence to error, not a controlled hard-easy experiment; full text could not be retrieved to confirm hard-easy-specific numbers.)

  • Overconfidence among financial planners / professional forecasters · 2009

    Surveys of financial planners and investors document overconfidence expressed as overly narrow forecast confidence intervals and excessive trading; calibration of professionals tends to be worse on genuinely uncertain (hard) forecasts, consistent with the hard-easy pattern. Reported via secondary literature reviews rather than a single primary hard-easy field study.

Damping

Use representative, randomly sampled items/cases rather than hand-picked 'hard' ones; calibrate against tracked outcomes; widen confidence intervals on genuinely difficult forecasts and seek base rates. Because much of the apparent effect is statistical (scale-end, regression, sampling), the strongest countermeasure is better measurement (frequency formats, outcome feedback) rather than just 'try to be less confident.'

Gigerenzer et al. (1991) and Juslin et al. (2000) show representative sampling reduces or removes the effect; Klayman et al. (1999) show frequency/two-choice formats yield less overconfidence than interval formats.

Reading the trace in the wild

Watch for confidence that barely moves between routine and genuinely tough decisions: forecasts on hard problems given the same narrow certainty as easy ones (overconfidence), or hedging on near-certain calls (underconfidence). Red flag phrases like 'about 90% sure' applied uniformly regardless of how hard the question actually is. Before trusting it as a bias, check whether the hard items were cherry-picked or whether you're just seeing regression to the mean.

First measured by Sarah Lichtenstein, Baruch Fischhoff, 1977 — Do those who know more also know more about how much they know?.

Adjacent channels

File your own case

Open the same case on your own draft.

Paste a memo, a research draft, or a strategy argument. It is scored against all 175 cards, and the strongest two or three risks come back with the evidence quoted and one practical next check.

Open a case on your draft →