The Peak of Not Knowing: Why Psychology's Most Famous Effect May Be Just a Statistical Mirage
🎧 Listen to this article
Psychology · 2026-07-15
Fully AI-generated article (no prior review).
The Hook: An Effect Everyone Knows – and Almost No One Knows Correctly
Few ideas in pop psychology have spread as effortlessly as the Dunning-Kruger effect. The British comedian John Cleese once distilled it perfectly: "If you are really, really stupid, then it's impossible for you to know you are really, really stupid." The effect shows up in LinkedIn posts, in boardrooms, in political arguments, in comment sections. It even has an iconic curve: a line that climbs steeply to a "Peak of Mount Stupid," then plunges into the "Valley of Despair," and finally rises gently toward the "Plateau of Sustainability." That curve is everywhere – except it isn't from Dunning and Kruger at all. It is a later invention, retroactively pinned onto the original finding.
This is symptomatic. The Dunning-Kruger effect is at once one of the best-known and one of the most misunderstood findings in psychology. Almost everyone wields it as a convenient weapon: to explain why other people – the incompetent, the arrogant, the political opponents – fail to notice how little they know. And therein lies a fine irony that runs through this entire article: confident use of the Dunning-Kruger effect is itself a perfect example of overconfidence. Most people who cite it have read neither the original study nor the devastating methodological critique that has been gnawing at it since 2016.
This article takes you the full distance. First we look at what Dunning and Kruger actually measured in 1999 – not the meme, but the data. Then we take the claim apart with the tools Sven values: we ask whether the effect is even real, or whether it inevitably arises from two mundane statistical phenomena, even in pure random noise. We see how the same famous graph can be produced from randomized data, without a single real human being. And in the end, a surprisingly nuanced truth remains, one wiser than either extreme – the naive meme version and the triumphant refutation.
Part 1: What Dunning and Kruger Actually Did
Four Studies, One Title for the Ages
In 1999, the two Cornell psychologists Justin Kruger and David Dunning published in the Journal of Personality and Social Psychology a paper with one of the most memorable titles in the literature: "Unskilled and Unaware of It: How Difficulties in Recognizing One's Own Incompetence Lead to Inflated Self-Assessments" (vol. 77, issue 6, pp. 1121–1134). The central thesis: people with low competence in a domain suffer from a double burden. They are not only bad at the task itself – they additionally lack precisely the metacognitive ability that would be needed to recognize that they are bad. Incompetence masks itself.
The design was simple. In one of the studies, the researchers gave 45 students a 20-question logic test (derived from an LSAT prep book). They then asked participants to rate themselves in two ways: first, how many questions they thought they had answered correctly (an absolute estimate), and second, how they had performed compared to the other test-takers (a relative estimate in percentiles). Further studies repeated the pattern with tests of grammar and – of all things – humor, where the "correct" answer on the humor test was defined by the judgment of professional comedians.
The Finding: The Famous Scissors
When Dunning and Kruger split their subjects into four groups (quartiles) by test performance and plotted both actual and self-estimated performance, the iconic picture emerged: two lines splaying apart like an opened pair of scissors. Actual performance rose steeply from the bottom to the top quartile. Self-estimated performance, by contrast, was remarkably flat – nearly everyone, regardless of how good they actually were, placed themselves in the upper-middle range.
The concrete numbers are more instructive than any meme. The lowest-performing quarter scored on average 10 out of 20, the top-performing quarter 17 out of 20. Yet both groups estimated they had gotten about 14 right. In absolute terms, the weakest overestimated their raw score by around 20 percentage points, while the best underestimated their performance by about 15 percentage points.
It got even more dramatic in the percentile comparison. The weakest test-takers believed they had beaten 62 % of the others; the strongest estimated their rank at 68 %. So both placed themselves near the upper-middle. But by definition, someone in the bottom quarter beats, on average, only 12.5 % of the others. From the gap between "I beat 62 %" and "I really beat 12.5 %" arises a spectacular overestimation of nearly 50 percentage points – the number that makes the meme so seductive.
The Metacognitive Explanation
Dunning and Kruger read this pattern not as statistics but as psychology. Their explanation was elegant and disquieting at once: the very skills you need to perform a task well are the same skills you need to judge how well you performed it. Someone who doesn't master the rules of logic can neither answer the test questions correctly nor recognize that their answers are wrong. The tool for performance and the tool for self-assessment are the same tool. If it's missing, it's missing twice.
To support this, they added a fourth study: if you teach the weakest participants to solve the task better, their self-assessment improves at the same time – they now recognize, in hindsight, how bad they had been. Competence, the punchline goes, grants you not only better performance but also better insight into your own former incompetence. The paper became a sensational success, cited tens of thousands of times, awarded the satirical Ig Nobel Prize, and within a few years turned into cultural common knowledge. But the more famous a finding, the harder the aftershock – and that aftershock came.
Part 2: The First Crack – The Better-than-Average Phenomenon
Almost Everyone Thinks They Are Above Average
To understand the critique, you need to know a second, much older psychological finding hidden inside the Dunning-Kruger curve: the better-than-average effect (BTA), also called "illusory superiority." People tend to place themselves slightly above average on almost any remotely positive trait. The evidence is overwhelming and at times comical: roughly 93 % of U.S. drivers consider themselves better drivers than average; about 90 % of college instructors believe they teach above average. Purely logically, this is impossible – the average cannot be exceeded by 90 %.
This very BTA effect provides a simple alternative explanation for the flat self-assessment line in the Dunning-Kruger graph. If everyone – competent or not – places themselves a little above the middle by default, then the self-assessment line is necessarily flat and sits above the middle. For the weak test-takers, it then gapes far apart from their (low) real performance; for the strong, hardly at all. The "scissors" opens – but not because the incompetent are especially clueless about themselves, but because virtually every human makes the same slight self-enhancement, and that enhancement deviates most from reality precisely for those who are already weak.
The crucial point: this explanation requires no special "double burden of incompetence." It requires only a single tendency toward self-enhancement, present equally in everyone – plus the arithmetic of extremes.
Part 3: The Second Crack – Regression to the Mean and the Miracle of Random Data
Why Extreme Values Almost Always Return Toward the Middle
Now comes the statistical core, and it is worth understanding cleanly, because it is the real explosive charge under the effect. Regression to the mean is an unavoidable mathematical phenomenon that occurs whenever two measures are not perfectly correlated – that is, essentially always, whenever measurement noise is involved.
The principle: take a person who scored an extremely low value on the first measurement. Part of that extreme value is "true" ability, part of it is bad luck (noise, a bad day, unlucky questions). On a second, independent measurement – say, self-assessment – the random component averages out. So the second measurement comes out, on average, less extreme, closer to the middle. Whoever was at the very bottom looks "better" on the second measure; whoever was at the very top, "worse." This is not a psychological mechanism but pure probability theory.
And precisely this pattern – bottom overestimates, top underestimates – is the Dunning-Kruger scissors. The question forces itself upon us: how much of the effect is psychology, and how much is merely regression plus better-than-average?
The Devastating Test: The Same Curve from Pure Chance
The answer came from critics with an elegant and destructive experiment. The mathematician Eric Gaze and colleagues generated 1,154 fictional "people" on a computer. Each was assigned a completely random test score – and an equally completely random self-assessment between 1 and 100. There was no psyche here, no ego, no metacognitive burden. Only random numbers.
Then they applied exactly the Dunning-Kruger procedure: they sorted the fantasy people into quartiles by test score and plotted the average self-assessment per quartile. The result was striking: the same famous scissors appeared. Because the random self-assessments averaged 50 in each quartile, while the bottom quartile really beats only 12.5 % of the others, pure noise produced an apparent overestimation of around 37.5 percentage points – without a single real human being. The effect, the cutting conclusion goes, is partly an artifact of research design, not of human thinking.
The Autocorrelation Critique
A related, still more fundamental critique was formulated pointedly by the economist Blair Fix in 2022, under the title "The Dunning-Kruger Effect is Autocorrelation." His argument: the classic Dunning-Kruger graph plots test performance on one axis and the difference between self-assessment and test performance on the other. This makes test performance appear on both sides of the equation – once on its own, and once (with a negative sign) inside the difference. You are effectively correlating a quantity with itself. Such "autocorrelation" necessarily produces the descending pattern that looks like the Dunning-Kruger effect – even when test performance and self-assessment are in truth completely independent, i.e., pure noise. Fix demonstrated this too, by reproducing the scissors from random data. I am of the opinion that this autocorrelation critique is the methodologically sharpest argument in the whole debate, because it does not merely name a possible confound but shows that the usual analytical procedure forces the effect by construction.
Part 4: The Clean Reassessment – Gignac & Zajenkowski 2020
How to Test the Effect Properly
Critique that only shouts "this is an artifact" remains unsatisfying as long as it doesn't say what a valid test would look like. That gap was filled by Gilles Gignac and Marcin Zajenkowski in 2020 in the journal Intelligence, with a paper whose title already gives away the verdict: "The Dunning-Kruger effect is (mostly) a statistical artefact." Their contribution is twofold: they expose why the usual quartile graphs are confounded (methodologically biased), and they propose statistically clean alternatives.
Two tools are central. First, the Glejser test for heteroscedasticity: the actual Dunning-Kruger hypothesis claims, after all, that the spread of self-assessment errors is larger among the incompetent than among the competent (the weak are further and more one-sidedly off). That is a statement about unequal error variance – and there are established tests for exactly that, without having to cut people into arbitrary quartiles. Second, nonlinear regression: you test directly whether the relationship between actual and self-estimated ability shows a curvature, as the theory demands.
The Result: A Linear Relationship, No Kink
On a sample in which objectively measured intelligence (a real IQ test) was compared with self-estimated intelligence, Gignac and Zajenkowski found: the relationship is essentially linear, with a modest positive correlation. The degree to which people misjudged their measured intelligence was roughly equal across the entire ability spectrum. There was no evidence that the incapable, specifically, are dramatically worse at self-assessment than the capable – the crucial, specific claim of the Dunning-Kruger effect simply did not appear in the data. What remained was the harmless, long-known finding: self-assessment correlates weakly, but positively, with actual ability.
The Dispute Continues: An Honest Debate
Science is not a triumphal march but a back-and-forth – and in the interest of honesty, that includes the critics themselves facing headwind. In a comment (from Avram Hiller, among others), it was objected that Gignac and Zajenkowski's finding of uniform error variance depended partly on a recoding decision: they mapped self-estimated relative intelligence onto a linear IQ scale, where a different scaling might have been more appropriate – and that choice could have shaped the result. The question of whether the Glejser test is really the right operationalization is also still being discussed. I am of the opinion that this ongoing dispute is no flaw but the hallmark of the debate: unlike the meme, which merely asserts, both sides here wrestle with the question of how to fairly test such a hypothesis in the first place.
Part 5: Can the Incapable Assess Themselves? The Direct Test
Nuhfer's Numeracy Studies
All the arguments so far have been indirect: they showed that the scissors arises even without psychology. But a direct, empirically answerable question remains: are the weakest test-takers really as hopeless at assessing themselves as the meme claims? This question was taken up by Edward Nuhfer and colleagues in two careful papers (2016 and 2017) in the journal Numeracy.
Nuhfer's team had students work through a 25-question test of scientific literacy and, after each question, rate how confident they were ("nailed it," "not sure," "no idea"). This allows a far more fine-grained analysis than the coarse quartile averages. The result clearly contradicts the meme: among the weakest students (the bottom quarter), only 16.5 % significantly overestimated their ability, while 3.9 % significantly underestimated it. That means: nearly 80 % of the "incapable" students judged their true ability fairly accurately. Little remains of the blanket idea that "the incompetent have no clue about their incompetence." Most people, in Nuhfer's conclusion, possess a perfectly functional sense of their own competence.
Part 6: What Remains – The Nuanced Truth
Time for an honest reckoning. It would itself be a Dunning-Kruger mistake to conclude hastily from the critique that "the whole effect is made up." The truth lies layered in between, and three statements can be held with a clear conscience.
First: the meme is wrong. The notion that the incompetent, specifically, play in their own league of reality-blindness and grope wildly off the mark, while the rest of humanity assesses itself soberly, is not supported by the data. A large part of the famous curve arises mechanically from regression to the mean, the better-than-average effect, and the autocorrelation in the analysis design.
Second: a real core remains. Even the sharpest critics concede that the accuracy of self-assessment correlates positively with actual ability – more competent people, on average, judge themselves somewhat more accurately. The signal is real; it is only smaller and far less dramatic than the meme suggests, and it holds gradually for everyone, not as a special curse of the incapable.
Third: the actual finding was always the better-than-average phenomenon. What Dunning and Kruger robustly showed is that most people consider themselves above average. That is a real, important, much-replicated result about human self-perception – just not the special, spectacular thesis that made their name famous.
For orientation, a comparison of myth and state of research:
| Aspect | The meme version | What the research suggests |
|---|---|---|
| Who overestimates themselves? | Only the incompetent, dramatically | Almost everyone, slightly (better-than-average) |
| The "mountain-valley-plateau" curve | Established by Dunning & Kruger | Later invention, not in the original |
| Cause of the scissors | Psychological "double burden" | Largely regression + BTA + autocorrelation |
| Are the weak blind to their ability? | Yes, utterly clueless | No, ~80 % judge themselves passably (Nuhfer) |
| Is there no effect at all? | — | There is: a weak but real positive correlation |
| The solid core | "Incompetence masks itself" | "People think they are above average" |
The Central Takeaway: Never Measure a Difference Against One of Its Own Components
The practical payoff of this article is by no means confined to psychology – it is a methodological warning that touches everyone who works with data, and thus every engineer, analyst, or developer. The deepest lesson of the Dunning-Kruger debate is this: whenever you plot a difference (prediction minus reality, estimate minus measurement, error) against a quantity that is itself contained in that difference, you will almost inevitably see a pattern – even when in truth there is none. Autocorrelation and regression to the mean are not exotic edge cases; they lurk in benchmark evaluations, A/B tests, performance reviews, model calibration plots, and in every "before-and-after" analysis.
Concretely, three moves that keep you out of the trap: (1) When you analyze error or improvement, never plot it solely against the baseline value without accounting for regression to the mean – instead compare against an independent second measurement or a random baseline. (2) Simulate your analysis procedure once with pure random data. If it already produces your "result" even then, the result is an artifact of the method, not of the world. This single move – reproducing the scissors from noise – is exactly what shook psychology's most famous effect. (3) Stay humble toward findings that fit your worldview too well. The Dunning-Kruger effect became so popular not least because it lets us regard others as clueless. A finding that flatters us so conveniently deserves double the skepticism.
Cross-References in the Vault
This article touches several topics already covered in this vault. The question of why effortless learning feels so deceptively competent is the direct sibling of the Dunning-Kruger question – both are about the illusion of competence (see Desirable Difficulties: Why Effortless Learning Deceives and Memory Lives on Effort). The epistemological base question of what it even means to know something – and how fragile the justification of beliefs is – is unfolded in the Gettier article (Three Pages Against 2,000 Years: The Gettier Problem and What Knowledge Really Is). And that our brain does not mirror perception but actively constructs it, systematically feigning certainty, is shown by the look at the predictive brain (The Predictive Brain: Predictive Processing and the Illusion of Perception).
A Closing Question for Reflection
In which area of your own skill would you be most likely to fall into the better-than-average trap – and how exactly do you know that your assessment there is correct, when, according to the metacognitive thesis, that is precisely where you would lack the very judgment tool needed to see your own gap?
Sources
- Kruger, J., & Dunning, D. (1999). Unskilled and Unaware of It: How Difficulties in Recognizing One's Own Incompetence Lead to Inflated Self-Assessments. Journal of Personality and Social Psychology, 77(6), 1121–1134. https://doi.org/10.1037/0022-3514.77.6.1121
- Gignac, G. E., & Zajenkowski, M. (2020). The Dunning-Kruger effect is (mostly) a statistical artefact: Valid approaches to testing the hypothesis with individual differences data. Intelligence, 80, 101449. https://doi.org/10.1016/j.intell.2020.101449 (PDF: https://gwern.net/doc/iq/2020-gignac.pdf)
- Nuhfer, E., Cogan, C., Fleisher, S., Gaze, E., & Wirth, K. (2016 & 2017). Random Number Simulations Reveal How Random Noise Affects the Measurements and Graphical Portrayals of Self-Assessed Competency and follow-up. Numeracy, 9(1) & 10(1). https://doi.org/10.5038/1936-4660.9.1.4 · https://doi.org/10.5038/1936-4660.10.1.4
- Gaze, E. C. / The Conversation (2023). The Dunning-Kruger Effect Isn't What You Think It Is. Scientific American / The Conversation. https://www.scientificamerican.com/article/the-dunning-kruger-effect-isnt-what-you-think-it-is/
- Fix, B. (2022). The Dunning-Kruger Effect is Autocorrelation. Economics from the Top Down. https://economicsfromthetopdown.com/2022/04/08/the-dunning-kruger-effect-is-autocorrelation/
- Hiller, A. Comment on Gignac and Zajenkowski, "The Dunning-Kruger effect is (mostly) a statistical artefact." Intelligence (2023). https://www.sciencedirect.com/science/article/abs/pii/S0160289623000132
- Overview: Dunning–Kruger effect. Wikipedia. https://en.wikipedia.org/wiki/Dunning%E2%80%93Kruger_effect
Note on evidence and reliability: The studies and critiques referenced (Kruger & Dunning 1999; Nuhfer et al. 2016/2017; Gignac & Zajenkowski 2020; Fix 2022) are peer-reviewed or publicly and verifiably documented, and were compiled for this article from primary and reputable secondary sources. Where I judge beyond the secured state of research, this is marked with "I am of the opinion that …".