Sven Erik Matzen

Software Architect | Cloud & Security Expert | AI-enabled Solutions

The 38 Witnesses Who Never Were: The Bystander Effect Between Laboratory Finding and Video Evidence

🎧 Listen to this article

Psychology · 2026-09-29

EU label: fully AI-generated content Fully AI-generated article (no prior review).

The Hook: A Headline That Shaped a Discipline

On March 27, 1964, The New York Times published an article by Martin Gansberg under the headline "37 Who Saw Murder Didn't Call the Police." It reported the killing of Catherine "Kitty" Genovese, a 28-year-old bar manager who had been attacked and murdered two weeks earlier outside her apartment building in Kew Gardens, Queens, by Winston Moseley. The point of the piece was not the murder. The point was the audience: dozens of neighbors had supposedly watched for more than half an hour, and no one had called the police.

That article became one of the most consequential newspaper stories in the history of the social sciences. A. M. Rosenthal, then the paper's metropolitan editor, turned it into a book — Thirty-Eight Witnesses. Two young social psychologists, John Darley and Bibb Latané, read the story, found the prevailing explanation ("big-city apathy," "alienation") unsatisfying, and asked a different question: what if the problem is not the indifference of individuals but their number? Four years later they had an answer that has appeared in virtually every introductory psychology textbook ever since.

And then something happened that makes this case fascinating for anyone who wants to understand how knowledge is produced and how it corrects itself: nearly every element of the original story turned out to be wrong. The number of witnesses was inflated, the extent of what they saw was exaggerated, and intervention did in fact occur. In 2016 the newspaper itself called its own reporting flawed, conceding that it had grossly exaggerated the number of witnesses and what they had perceived.

The obvious conclusion would be: the myth has been debunked, therefore the effect has been debunked. That conclusion is wrong, and instructively so. The bystander effect is one of the best-replicated findings in social psychology — more than 100 independent effect sizes measured across five decades, with a clear average effect. At the same time, analyses of real surveillance footage from three countries show that in nine out of ten public conflicts at least one person present intervenes — and that more bystanders do not lower the chance of help but slightly raise it.

Both findings are solid. Both have been reproduced. And they do not contradict each other. Seeing why requires looking very closely at what was actually measured. This article reconstructs the case, the finding, the counter-finding — and the algebra that reconciles them. At the end there is a practical consequence familiar to anyone who has ever posted an alert to a group channel instead of to a person.


Part 1: What Actually Happened in 1964

The Story as It Was Told

The canonical version runs like this: 38 people watched a 35-minute attack in three phases from their windows. They saw the stabbing. They heard the screams. No one called the police. No one went downstairs.

This version has enormous narrative power, because it poses a moral puzzle that cannot be explained by malice. Thirty-eight evil people in one apartment block is implausible. Thirty-eight ordinary people doing something monstrous is a problem that demands a structural explanation.

What the Trial Record Shows

In 2007, Richard Manning, Mark Levine, and Alan Collins examined the factual basis of the story in the American Psychologist, working from the court record of Moseley's trial. Their findings dismantle the narrative on three counts.

First, the number. Five residents of the building testified at trial; only three were eyewitnesses who saw Genovese together with her attacker. The assistant district attorney on the case later stated that only about half a dozen people had been found who saw what was going on. That more people heard screams is not disputed — but hearing is not seeing, and a scream at night in a neighborhood with a bar on the corner was an ambiguous signal in 1964.

Second, what they saw. The statements of the three eyewitnesses describe unspecific scenes. One saw a couple standing close together, not fighting. Another, Robert Mozer, saw a woman kneeling on the ground. None of the witnesses reported having seen the stabbing — and that is not self-exculpation, it is geometry: after the first attack Genovese walked around the corner of the building, out of sight of the windows. The second, fatal attack took place in an entryway that no one could see into from the apartments.

Third, the inaction. People did intervene. Robert Mozer shouted from his window and frightened the attacker off — which is precisely why there was the interruption that makes the story dramatic in the first place. And Sophia Farrar, a neighbor, after being telephoned by another resident, ran into the entryway and held the dying Genovese in her arms, without knowing whether the killer was still there. Anyone looking for an unarmed woman who walked into a situation whose danger she could not assess will find her in this story — she just did not appear in the headline.

Accounts of the calls to the police remain contradictory to this day. It is plausible that at least one was made; a complicating factor was that in 1964 New York still had no unified emergency number.

Why the Myth Was Nonetheless Productive

Manning, Levine, and Collins go further, and this is the more interesting part of their argument: the parable did not merely transmit a false picture, it narrowed the research agenda. It established a particular image of the collective — the group as an agency that imposes inaction on the individual. In doing so it displaced the older research tradition, which had viewed groups primarily as prone to action and violence, and for decades it obscured the obvious complementary question: under what conditions do groups facilitate intervention?

That question was only asked once video material became available. And it produced the findings in Part 4.

It is worth pausing here. The story of the 38 witnesses is a textbook case of how a compelling narrative takes on the role of a data point — the same mechanism described at the level of individual memory in The Invented Memory: Why Recollection Is Not a Recording – and How Elizabeth Loftus Showed You Can Remember Things That Never Happened. A discipline can hold a false memory.


Part 2: The Laboratory Finding

The Seizure Experiment

In 1968 Darley and Latané published, in the Journal of Personality and Social Psychology, the experiment that gave the effect its name. The design was chosen to isolate the suspected cause and rule out everything else.

Students took part in a purported group discussion about problems of university life. To avoid embarrassment, the explanation ran, each participant would sit alone in a cubicle; communication would run over intercoms, two minutes per person, in fixed rotation, with no experimenter present. In fact the only real participant was the subject; all the other voices came from a tape.

On the second round, one of the voices ran into trouble: halting speech, then the unmistakable description of an oncoming epileptic seizure, then choking, then silence. What was measured was whether and how quickly the subject left the cubicle to get help.

The only experimentally varied quantity was how many other participants the subject believed were listening:

Perceived group size Proportion who intervened Character of response time
2 people (victim and subject only) 85% fast, around one minute on average
3 people 62% delayed
6 people 31% strongly delayed, well over two minutes on average

The finding is strong for two reasons. First, the independent variable is only the believed presence of others — nobody saw a passive counterpart whose calm could have been contagious. Second, help declines not just in frequency but in speed, and both vary monotonically with group size. That is a dose-response pattern, not a fluke.

One detail that usually gets lost in textbook summaries deserves attention: in the debriefing, the non-helpers did not come across as indifferent but as agitated. They had not decided not to help. They were stuck in a conflict and could not get out of it. That observation becomes important in Part 5.

The Room That Filled With Smoke

In the same year, Latané and Darley published a second study with a different mechanism. Participants filled out a questionnaire in a waiting room while white smoke streamed in through a vent in the wall — for up to four minutes, by the end thick enough to impair visibility.

Alone in the room, 75% reported the smoke within the first two minutes. With two confederates present who conspicuously ignored it, only 10% reported. In groups of three genuine, uninstructed subjects, only 38% of the groups produced a reporter at all.

What operates here is not diffusion of responsibility but a second mechanism: pluralistic ignorance. The situation is ambiguous. Everyone looks to the others for an interpretation. All the others are doing exactly the same thing at that moment — glancing around discreetly and staying calm so as not to embarrass themselves. The result is a group of people mutually reassuring each other through their own façades. Every individual is uneasy, and every individual believes they are the only one who is.

A year later Latané and Judith Rodin completed the picture with the "lady in distress" design — an experimenter audibly falls from a chair in the next room. Alone, about 70% helped; in the presence of a passive stranger, 7%; among pairs of subjects who were strangers to each other, at least one intervened in 40% of pairs — among pairs of friends, around 70% again. Friends read each other better and fear humiliation in front of each other less. Inhibition, then, is not tied to the mere presence of bodies but to the social relationship between them.

The Three Mechanisms

This series produced the explanatory framework still in use today. There are three separate processes that happen to point in the same direction:

  1. Diffusion of responsibility. The perceived personal obligation to act is divided among everyone present. With six people, each feels responsible for a sixth. This mechanism works even when you cannot see the others — which is why the cubicle experiment worked.
  2. Pluralistic ignorance. The situation is ambiguous, and the others' calm is read as evidence that nothing is wrong. This mechanism requires line of sight and disappears as soon as the situation becomes unambiguous.
  3. Evaluation apprehension. Intervening means acting in front of an audience and possibly embarrassing yourself — a false alarm, clumsy first aid, a fight that turns out to be a couple's quarrel. This mechanism can also increase helping when the norm of intervening is clear and the audience is disposed accordingly.

In 1981, after ten years of research, Bibb Latané and Steve Nida drew a remarkably unambiguous conclusion in the Psychological Bulletin: the group-size effect was among the most robust findings in the field. For the rest of the twentieth century, the matter was considered settled.


Part 3: How Strong Is the Effect Really?

The 2011 Meta-Analysis

In 2011 Peter Fischer and colleagues published in the Psychological Bulletin the most comprehensive quantitative synthesis to date: 105 independent effect sizes from studies spanning 1960 to 2010, together involving more than 7,700 subjects. The overall result: g = −0.35.

That number is first of all a confirmation. A medium negative effect across five decades, laboratories, countries, and paradigms is a strong result in social psychology — especially in a discipline in which the replication crisis has, since the 2010s, qualified numerous textbook findings, as described in The Marshmallow Promise: How a Candy Test Redefined Willpower for Half a Century – and Was Then Put to the Test Itself and The Peak of Not Knowing: Why Psychology's Most Famous Effect May Be Just a Statistical Mirage. The bystander effect does not belong in that category. It is real.

But it is also: medium-sized, not overwhelming. And above all it is highly conditional. The moderator analysis is the genuinely valuable part of the paper, because it tells us when the effect becomes small or vanishes. Inhibition was substantially reduced when

  • the situation was perceived as dangerous rather than harmless,
  • a perpetrator was present rather than the event being a mere accident,
  • helping required physical rather than merely informational involvement,
  • the bystanders were exclusively male,
  • the others present were naive fellow subjects rather than instructed passive confederates,
  • the others were only virtually present rather than physically present,
  • the people involved were acquainted rather than strangers.

Fischer and colleagues explain this with the arousal-cost-reward model: an unambiguously dangerous situation is recognized faster and more reliably as a real emergency, produces higher arousal, and therefore more helping. At the same time, in dangerous situations the presence of others provides a benefit it does not provide in harmless ones — confronting an attacker two-against-one is objectively less risky than alone.

The Crucial Qualification

This puts the classic formulation of the effect off balance. In its popular form it read: the more onlookers, the less likely help becomes. The meta-analysis says, more precisely: the more onlookers, the less likely it is that a given individual helps — and this relationship weakens or reverses the more unambiguous and dangerous the situation is.

That is not the same statement, and the difference between "a given individual" and "anyone at all" is the core of the next chapter. But first we need data from the real world. Because almost everything cited so far comes from staged situations — for good reasons, but with an obvious question in the background: do people in a real street conflict behave the way students in a cubicle do?


Part 4: The Video Evidence

219 Real Conflicts

In 2020 Richard Philpot, Lasse Suonperä Liebst, Mark Levine, Wim Bernasco, and Marie Rosenkrantz Lindegaard published a paper in the American Psychologist that bypassed the field's methodological bottleneck. Instead of staging emergencies, they analyzed recordings of real ones: 219 video clips of public conflicts from the surveillance camera systems of three inner cities — Amsterdam (63), Lancaster (95), and Cape Town (61).

The choice of cities is not arbitrary. They differ considerably in levels of violence and in the subjective sense of safety; Cape Town is regarded as markedly less safe than Lancaster. If intervening were a function of perceived safety, that should show up.

The results:

Finding Value
Conflicts with at least one intervening bystander 90.9%
Average number of interveners per incident 3.76 (SD = 3.01)
Effect of each additional bystander on the chance of help OR = 1.10 (p = .008)
Differences between the three cities not significant

In nine out of ten cases at least one person present did something — stepping in, separating the parties, restraining the aggressor, pulling the victim away, talking people down. And the number of onlookers did not lower the chance of help, it raised it slightly. The finding was stable across three very different urban contexts.

The Role of Danger

Two years later, Lindegaard, Liebst, Philpot, Levine, and Bernasco followed up in Social Psychological and Personality Science with the obvious next question: does intervening depend on the level of danger? They coded 80 interpersonal conflicts from the same three cities (Amsterdam 26, Lancaster 30, Cape Town 24; number of bystanders ranging from 1 to 62, median 15) into three levels of aggression: intrusive but non-violent acts (pointing, light pushes); physical aggression (punches, kicks, shoves); and weapon use or violence against a person already kneeling or lying on the ground.

The finding has two halves. The odds of intervention were roughly 19 times higher as soon as targeted aggression became visible at all. A further increase in danger — from level 1 to 2 to 3 — by contrast did not significantly change the likelihood of intervention.

That is a threshold effect, not a gradient. People do not intervene in proportion to danger, but as soon as a situation crosses the line from "unclear" to "unmistakably an assault." And that is precisely the empirical confirmation of the mechanism that Latané's smoke experiment had suggested more than fifty years earlier: the decisive variable is not courage but unambiguity.

What These Studies Do Not Show

Honesty requires three qualifications. First: surveillance cameras stand in busy inner-city locations — nightlife districts, squares, station forecourts. That is a specific slice of reality, not a cross-section. A back alley at night is not part of it.

Second: measurement was at the event level ("did anyone intervene?"), not at the individual level ("did this particular person intervene?"). With a median of 15 bystanders and an average of 3.76 interveners, the large majority of those present remain passive. "Help was given" does not entail "most people helped."

Third: what counted as intervention was deliberately defined broadly. Talking someone down is a different thing from placing yourself between a knife and a victim. The studies say that inaction is not the norm. They say nothing about whether the help was effective.


Part 5: Why Both Are Right — the Algebra Behind It

Two Questions That Sound Like One

The apparent contradiction between laboratory and video dissolves entirely once two quantities are separated that ordinary language runs together:

  • p — the probability that a given bystander intervenes.
  • P — the probability that at least one of the bystanders intervenes.

If the decisions were independent, then

$$P = 1 - (1 - p)^n$$

with n as the number of people present. The bystander effect is a statement about p. The video studies measure P. And p can fall with n while P nevertheless rises — as long as p falls more slowly than 1/n.

A worked example makes this tangible. Suppose a single bystander alone intervenes with probability 60%, and this individual readiness drops sharply as the group grows:

Bystanders n individual readiness p chance of help \(P = 1-(1-p)^n\)
1 0.60 60.0%
2 0.45 69.8%
5 0.30 83.2%
10 0.22 91.7%
20 0.15 96.1%

Individual readiness has collapsed from 60% to 15% — a dramatic bystander effect. And yet the chance that help is given rises from 60% to 96%. Both sentences describe the same table: "each individual becomes markedly more passive" and "help becomes far more likely."

Anyone who recognizes the pattern is right: it is exactly the same algebra that, in The Long Tail of Latency: Tail Latency and Why Averages Lie in the Cloud, explains why with 100 subrequests at a 1% outlier rate each, 63% of user requests turn slow. There, fan-out is the problem; here it is the rescue — the formula is identical, only the sign of what we want is reversed.

This also makes the OR = 1.10 per additional bystander in Philpot et al. interpretable. It is precisely what this algebra predicts when p declines moderately — or when, in a public, unambiguously aggressive conflict, the inhibitory component is weak to begin with, as Fischer's moderators suggest.

The Second Resolution: Reflex Before Decision

There is a second, independent resolution. In 2018 Ruud Hortensius and Beatrice de Gelder proposed, in Current Directions in Psychological Science, that the bystander effect be understood not as the outcome of a deliberation but as the outcome of a reflexive action system.

Their starting observation is the one from Part 2: non-helpers are not indifferent but aroused. That fits poorly with a model in which someone rationally divides responsibility by six. It fits well with a model of two opposing motivational systems:

  • System I — personal distress: an immediate, self-oriented response that prepares freezing, flight, or fight. It is amplified by the presence of others.
  • System II — sympathy: a slower, other-oriented response that supports helping.

Whether help occurs is the resultant of both systems, modulated by personality traits. Neuroimaging findings support this: in an fMRI study by the authors, activity decreased with a growing number of bystanders precisely in the regions relevant to preparing a helping action — the pre- and postcentral gyrus and the medial prefrontal cortex. It is not that the decision is made differently; the preparation to act is dampened before a decision even arises.

I am of the opinion that this model also explains why moral appeals accomplish so little and training so much: you cannot change a reflex by persuasion, but you can rebuild it through repetition. This reading goes beyond the evidence and should be taken as an interpretation, not a result.


Part 6: The Digital Arena

The question of how all of this behaves in online environments is practically relevant — for community moderation, for reporting mechanisms, for internal communication platforms.

Fischer's meta-analysis already contained a hint: the inhibitory effect was weaker when the others were only virtually present. That makes sense, because two of the three mechanisms depend on co-presence. Pluralistic ignorance requires the observable calm of others; evaluation apprehension requires an audience that will witness your embarrassment. In a chat with 400 members but nobody visibly watching, what mainly remains is diffusion of responsibility — and that is the mechanism that only has to believe others are present.

In 2018 Franccesca Kazerooni, Samuel Hardman Taylor, Natalya Bazarova, and Janis Whitlock examined, in the Journal of Computer-Mediated Communication, a variable that has no clean offline equivalent: not the number of onlookers but the number of offenders, and whether an attack is original or reshared. Their finding: four offenders posting their own content significantly increased perceived hurt, the classification of the event as cyberbullying, and the willingness to intervene, compared with a single offender. When the same content was reshared rather than originally written, however, willingness to intervene declined noticeably.

This is consistent with the unambiguity finding from Part 4: four independent attacks turn an ambiguous incident into a recognizable assault. And it adds an uncomfortable digital peculiarity: resharing appears to dilute not only the responsibility of the onlookers but also that of the attacker — and with it the onlookers' willingness to treat anything as an assault at all.


Part 7: The Overarching Principle

Step back, and this case tells two stories layered on top of each other.

The first is about scientific self-correction, and it does not unfold the way the cliché suggests. The original story was largely false. The finding it triggered is largely correct. The popular formulation of that finding was too coarse. And the correction of that formulation did not come from a failed replication attempt but from a new data source that could finally pose an old question at the right level. Anyone who infers from the fact that the Genovese story is a myth that the bystander effect is a myth is making the same mistake in reverse: letting the narrative stand in for the evidence once again.

The second story is about levels of measurement. Almost every apparent contradiction between two solid findings is in truth a confusion of the unit of analysis. Individual or event. Request or subrequest. User or session. Commit or release. The question "a rate of what, exactly?" resolves more debates than any additional study could.

The Central Takeaway

The practical lesson: the bystander effect is not a character deficiency but an addressing problem — and you fix it not through appeals but by removing ambiguity and assigning responsibility by name.

Both levers follow directly from the evidence. Unambiguity, because the odds of intervention jump by a factor of 19 as soon as an incident becomes recognizable as one — but do not rise further after that. Naming, because diffusion of responsibility is the only one of the three mechanisms that works without line of sight, and therefore the only one that continues unchecked in distributed and digital environments.

Four concrete steps follow:

  1. In an emergency, as the person affected: address individuals. Not "Help!" but "You, in the blue jacket — call emergency services." That sets n to 1 for that person. The same applies if you arrive first and find an audience: hand out tasks with eye contact and a description, do not shout into the crowd.
  2. In an emergency, as a bystander: go first. The first intervention changes the situation for everyone else twice over — it removes the ambiguity and establishes a visible norm. The 3.76 interveners per incident in the video data are almost never 3.76 independent decisions; they are a cascade.
  3. In organizations: no alert without a name. An error report in a channel with 200 members is the smoke experiment with a membership list. An escalation needs a named person with a time window — and a rule for what happens if that person does not respond. That distributed systems must solve exactly this problem explicitly — who determines that something has failed, and who acts on it — is described in The Whisper of Machines: Gossip Protocols, SWIM, and How a Cluster Learns Who Is Still Alive. Humans are rarely handed that explicitness for free.
  4. In code review and quality assurance: ownership before audience. "Everyone take a look" reliably produces less attention than two reviewers assigned by name. This is not a metaphor but the same mechanics in an environment where the cost of not looking simply arrives later.

A Question to Reflect On

The video data say that in nine out of ten public conflicts help is given, and the laboratory findings say that each individual in a crowd becomes markedly more passive. Both are true because enough people are present that the willing few suffice. The statistics carry the outcome, not the disposition of the individuals.

From this follows an uncomfortable question: our systems — emergency chains, reporting paths, on-call rotations, review processes — often work not because their responsibilities are cleanly defined, but because enough people are attached that statistically someone almost always turns up. Which of those chains in your own environment would still hold if only three people were watching instead of fifteen — and how do you know, without having tried it?


Cross-References in the Vault


Sources

  • Darley, J. M., & Latané, B. (1968): Bystander intervention in emergencies: Diffusion of responsibility. Journal of Personality and Social Psychology, 8(4), 377–383. DOI: 10.1037/h0025589
  • Latané, B., & Darley, J. M. (1968): Group inhibition of bystander intervention in emergencies. Journal of Personality and Social Psychology, 10(3), 215–221. DOI: 10.1037/h0026570
  • Latané, B., & Rodin, J. (1969): A lady in distress: Inhibiting effects of friends and strangers on bystander intervention. Journal of Experimental Social Psychology, 5(2), 189–202. DOI: 10.1016/0022-1031(69)90046-8
  • Latané, B., & Nida, S. (1981): Ten years of research on group size and helping. Psychological Bulletin, 89(2), 308–324. DOI: 10.1037/0033-2909.89.2.308
  • Manning, R., Levine, M., & Collins, A. (2007): The Kitty Genovese murder and the social psychology of helping: The parable of the 38 witnesses. American Psychologist, 62(6), 555–562. DOI: 10.1037/0003-066X.62.6.555
  • Fischer, P., Krueger, J. I., Greitemeyer, T., Vogrincic, C., Kastenmüller, A., Frey, D., Heene, M., Wicher, M., & Kainbacher, M. (2011): The bystander-effect: A meta-analytic review on bystander intervention in dangerous and non-dangerous emergencies. Psychological Bulletin, 137(4), 517–537. DOI: 10.1037/a0023304
  • Hortensius, R., & de Gelder, B. (2018): From empathy to apathy: The bystander effect revisited. Current Directions in Psychological Science, 27(4), 249–256. DOI: 10.1177/0963721417749653
  • Kazerooni, F., Taylor, S. H., Bazarova, N. N., & Whitlock, J. (2018): Cyberbullying bystander intervention: The number of offenders and retweeting predict likelihood of helping a cyberbullying victim. Journal of Computer-Mediated Communication, 23(3), 146–162. DOI: 10.1093/jcmc/zmy005
  • Philpot, R., Liebst, L. S., Levine, M., Bernasco, W., & Lindegaard, M. R. (2020): Would I be helped? Cross-national CCTV footage shows that intervention is the norm in public conflicts. American Psychologist, 75(1), 66–75. DOI: 10.1037/amp0000469
  • Lindegaard, M. R., Liebst, L. S., Philpot, R., Levine, M., & Bernasco, W. (2022): Does danger level affect bystander intervention in real-life conflicts? Evidence from CCTV footage. Social Psychological and Personality Science, 13(4), 795–802. DOI: 10.1177/19485506211042683

← All articles