Everybody knows you do not do that. Nobody thinks that is okay. Obviously that is inappropriate.
We use sentences like these constantly. They sound like reports about society, but they are often something else wearing a fake mustache: a personal judgment promoted to the position of public consensus.
A new study published Sept. 1 in Communications AI & Computing offers an unusually clean look at that promotion. Researchers asked six large language models and 320 Americans to estimate how people in the United States collectively judge the appropriateness of ordinary behaviors in ordinary settings. The models were much better than the average individual person at guessing the population average. The reason was not that the machines discovered an objective code of manners. It was that humans kept insisting the code was more absolute than it actually was.
The Robot Did Not Become Miss Manners
The study, led by Kimmo Eriksson and colleagues, tested 555 scenarios built by combining 37 everyday behaviors with 15 situations. Think argue, cry, laugh, read or pray audibly, placed into settings such as a bar, a church, a job interview or someone's own room. The target was not a philosopher's answer to whether the behavior was good. It was the average appropriateness rating previously given by a large U.S. sample on a scale from 0 to 9.
That distinction matters. The models were not asked, "Is this appropriate?" They were effectively asked, "What number do you think Americans, on average, would give this?" The human estimators received the same basic task. Each person estimated a random subset of 50 scenarios. Six LLMs - GPT-5.2, Claude Sonnet 4.5, Gemini 3 Flash, Llama 3.1 8B, Llama 3.1 70B and DeepSeek V3 - were queried 15 times on every scenario.
All six models beat the average human estimator. On the study's primary measure, mean absolute error, the average person missed the measured norm by 1.77 points on the 0-to-9 scale. The open-weight models landed between 1.22 and 1.62. Two proprietary models even beat the best individual human participant, whose error was 0.92.
That is the headline result and it's also the least interesting part.
Humans Kept Erasing the Middle
When the researchers looked at the shape of the answers, the human problem became visible. The actual U.S. norms contained plenty of middle-range judgments. People, by contrast, disproportionately predicted ratings near the endpoints. They tended to imagine that the population saw a behavior as basically right or basically wrong, even when the measured public judgment was more mixed.
The authors call this "categorical bias." People were explicitly asked to estimate an average, but many answered as though they were rendering a verdict. The more frequently a participant used the extreme ends of the scale, the worse that participant was at estimating the actual population average.
This is where the study wanders out of AI research and directly into ordinary human conversation. We are very comfortable replacing "I think this is inappropriate" with "people think this is inappropriate." We replace "this bothers me" with "nobody does that." We replace "my circle treats this as obvious" with "everybody knows." The grammar quietly converts a personal rule into a social fact.
That does not mean the personal judgment is wrong, it means the claim about everybody else may be imaginary.
"Normal" Is Usually a Distribution, Not a Switch
There is a useful conceptual correction hiding here. A social norm, at least as measured in this study, is not a light switch marked NORMAL and ABNORMAL. It is a distribution of judgments across people. The average can sit in the middle because a lot of people feel moderately about something, because the population is split, or because context introduces uncertainty. Those are different social realities, but none of them justify pretending there is unanimous agreement. The middle is not necessarily indecision because sometimes the middle is disagreement accurately measured.
Humans appear to have trouble preserving that distinction when asked to imagine the group. The researchers argue that the pattern is consistent with the false-consensus effect: we know our own reaction first, and then we overestimate how widely it is shared. If I experience a behavior as clearly rude, it is cognitively convenient to assume the room comes with me.
An LLM has a peculiar advantage on exactly this task. It has no childhood dinner table, no friend group, no church basement, no office clique and no neighbor whose habits have been irritating it since 2014. It has instead absorbed enormous quantities of language describing how people talk about situations. For estimating a population average, that broad statistical exposure may be more useful than one person's vivid but narrow social experience.
That does not make the model socially wiser. It makes it, in this narrow experiment, a better weather station for the center of the conversation.
Then the Humans Come Back
There is a second result that keeps this from becoming another story about machines replacing people. Individual humans were noisy, biased and frequently overconfident. Groups of humans were excellent.
The errors made by different people were mostly independent. One person overestimated one situation; another missed somewhere else. When the researchers averaged human judgments together, those idiosyncratic mistakes began canceling one another out. The mean absolute error fell from 1.77 for the average individual to 0.57 for a collective of 15 people - better than the best LLM result reported in the aggregation analysis.
The models had the opposite problem. Their answers were individually strong but their mistakes resembled one another. Across models, error correlations averaged 0.43, compared with just 0.06 between humans. There were 37 scenarios all six models overestimated and 59 they all underestimated. Asking the same model repeatedly did almost nothing, and combining all six models produced an error of 0.73, statistically no better than the best model alone at 0.71.
The machine knows a lot, but apparently the machines know a lot of the same things wrong. Humans, meanwhile, contain the one feature that looked inefficient at the individual level and became valuable at the group level: disagreement.
The Best Answer Was Neither Human Nor Machine
When the researchers averaged a 15-person human collective with the strongest LLM estimate, the error dropped to 0.48. That was a 32.5 percent reduction in error compared with the model alone and a 15.9 percent improvement over the human collective alone.
That result deserves more attention than the competition framing. The model was good at locating the broad center. The people supplied diverse errors the model did not share. Each corrected something about the other.
There is a larger lesson in there for the way AI is often discussed. We keep staging human intelligence and machine intelligence as if somebody must leave the building. This study instead found a case where the useful property of human judgment was not flawless expertise. It was heterogeneity. The people were valuable partly because they were not wrong in identical ways.
A crowd can be smart because its members disagree. An AI ensemble can fail to become a crowd because its members agree too much.
Before We Appoint the Robot Mayor of Normal
There are limits worth keeping attached to the finding. The study concerns estimates of average U.S. norms, not the ability to behave appropriately in a live room. The authors explicitly note that real social competence also requires reading a specific audience, adapting in real time and recognizing that a local group may differ from a national average.
The human samples also came from Prolific, and the authors caution that regional and community differences within the United States may not be represented proportionally. The study was not preregistered, although the human data were collected in two waves that produced nearly identical accuracy results.
And "American norms" are not "human norms." A separate 2025 study involving more than 25,000 participants across 90 societies found substantial cultural variation in everyday appropriateness judgments and evidence that some norms change over time. Normal has a place. Normal has a date. Normal has a sample.
So the result should not be translated into "AI knows what is right." That would repeat the exact categorical mistake the paper is warning us about.
Everybody Might Want to Sit This One Out
Every time somebody says, "Everybody knows," there are at least two claims hiding inside the sentence. The first is the speaker's judgment. The second is the assertion that the judgment is widely shared. We tend to treat those as one thing. They are not.
Sometimes you have a principle. Sometimes you have a preference. Sometimes you have a local norm. Sometimes you have a widely shared social expectation. Sometimes you have five people in your group chat who all agree with you. Those categories can feel identical from inside your own head.
They are not identical from the outside. The robot did not discover what normal means. It was simply less tempted, on this task, to turn the middle into an endpoint. And when actual people were combined, their disagreement reconstructed the middle even better.
That leaves us with an annoyingly human conclusion. The individual person may be too certain. The individual machine may be too uniform. The best estimate comes from keeping both the pattern and the disagreement in view.
So the next time a sentence begins with "Nobody thinks..." or "Everybody knows..." it may be worth checking whether everybody has, in fact, been consulted.
Apparently the most normal thing about normal is that we do not agree on it nearly as much as we think.
What the study actually measured
| Target | Average U.S. appropriateness ratings for 555 everyday behavior-in-context scenarios. |
|---|---|
| Human estimators | 320 U.S. adults; each estimated a random subset of 50 scenarios. |
| LLMs | Six models tested in January 2026, with 15 runs per scenario. |
| Key individual result | Average human MAE 1.77; every tested LLM performed better. |
| Crowd result | Aggregating 15 human estimates reduced MAE to 0.57. |
| Hybrid result | A 15-person human aggregate combined with Gemini 3 Flash reached MAE 0.48. |
| Important caveat | This benchmark measures estimation of a population average, not moral truth or real-time social competence. |
SOURCE NOTES
• Eriksson, Karlsson, Vartanova et al., "Large language models outperform humans at estimating society's everyday norms," Communications AI & Computing, Sept. 1, 2026
• Peer-review history for the 2026 paper, Communications AI & Computing
• Eriksson, Strimling, Vartanova et al., "Everyday norms have become more permissive over time and vary across cultures," Communications Psychology, Oct. 7, 2025
The 2026 study was not preregistered. Its authors report closely replicated human-estimation accuracy across two collection waves. The study is U.S.-specific and uses population-average appropriateness as its norm benchmark; it should not be read as showing that LLMs possess superior moral judgment or general social intelligence.