Chinese Whispers at Civilizational Scale
The truth is becoming irrecoverable
AI systems learn by absorbing enormous quantities of text, billions of pages of books, articles, websites, and social media posts. If we think of this as their “education,” then the quality of what they produce depends on the quality of what they learned from.
In 2024, researchers at Oxford, Cambridge, and the University of Toronto demonstrated what happens when the next generation of AI models learns not only from human-written text but from text that earlier models produced. They called it “model collapse,” which they described as a degenerative process in which, as the researchers put it:
“the model becomes poisoned with its own projection of reality.”
The things that were already common in the materials the model learned from get reinforced, while what was uncommon gradually fades. Each generation “remembers” less of what was unusual and more of what was already dominant. And over enough generations, what the AI produces bears less and less resemblance to the world it was supposed to represent, and the process, once underway, doesn’t reverse.
Over the past few years, I've been studying influence networks. I wanted to understand how topics and narratives evolve across social media. More specifically, I wanted to understand how distortions in truth and fact take shape, gain traction, and harden into things people believe.
The sheer volume of AI-generated content now circulating across platforms is being shared, amplified, and absorbed back into the material that the next generation of AI systems will learn from. This points to a problem that I think is different in kind from disinformation (as it's usually discussed), and in some ways harder to confront. It doesn't require anyone to be lying, at least not at the level of the mechanism itself. It's a feedback loop by which AI doesn't just spread false information but absorbs it, reproduces it, and over time makes the truth irrecoverable. And once this loop is running, it's very difficult to stop.
How the loop works
False narratives enter the information ecosystem through: social media platforms, news outlets, state-backed media, activist organizations, and even academic literature.
These channels vary in how much authority people assign to them.
The narratives get amplified as they circulate across these channels - being shared, reposted, retweeted, commented on, referred to, picked up by other outlets. In the course of this process they contaminate the vast pools of text that AI models learn from. Those then serve the distorted narratives back to users as authoritative, knowledgeable-sounding responses, which too get shared, discussed, and quoted and referended, further eroding the shared factual record.
Each pass through this loop compounds the distortion, and at some point this reaches a threshold where the ground truth becomes lost and inaccessable because the AI-generated layer of information has overwhelmed the primary sources which have become buried under layers of synthetic reproduction.
This feedback loop has an asymmetry that, I think, most people haven’t fully registered yet. What actually happened on October 7th - what people did, what they said, who killed whom - happened once. The footage exists or it doesn’t. The forensic report has been written or it hasn’t. The survivor testimony has been given or it hasn’t. You can’t go back and gather more evidence of something after the fact.
But AI-generated content is different. It multiplies. Every time OpenAI or Google or Anthropic builds the next version of its AI models using newer text from the web, the system doesn’t just inherit whatever distortion was in its training data, it produces new content carrying that distortion forward. And every sentence it produces sounds authoritative, because sounding authoritative is what the system was built to do.
So the truth is finite. The distortion compounds
Evidence and its echo
Consider two events with very different evidentiary profiles.
The October 7th massacre is extensively documented. Body-camera footage, satellite imagery, forensic reports, hostage testimony, and thousands of contemporaneous social media posts from both victims and perpetrators all form part of the record. The signal is so distinctive, and spread across so many different forms of evidence, that false narratives can't easily outweigh it in the text AI systems learn from. To do that, you'd need not just volume but a coherent counter-narrative that accounts for each evidentiary strand.
Contrast that with an event like the 2018 Douma chemical attack in Syria, where dozens of civilians were killed in what Western governments attributed to an Assad regime chlorine strike while Russia and Syria called it a fabrication by rebels. The primary evidence is a handful of soil samples, a few minutes of unverified video, and contested OPCW inspection reports from which inspectors who had been on the ground publicly dissented, claiming the final report didn’t reflect their findings. The leaked dissent itself became a vector for competing narratives.
When the evidentiary base is this shallow, and already this disputed, even a modest amount of AI-generated content reinforcing one version over another can shift the balance within just a few cycles of companies rebuilding their AI systems using newer text from the web.
You’d expect an event as extensively documented as October 7th to be far more resilient to this kind of distortion than event like what happened in Douma, and in one sense it is: the sheer volume and diversity of primary evidence makes it much harder for any false counter-narrative to account for everything. But this fact is misleading because most of that evidence doesn’t exist in the form that AI systems actually learn from.
Body-camera footage, forensic testimony, and court records sit in video archives, legal databases, and gated institutional repositories. They aren’t available as text on the open web. What the AI models encounter when they’re learning is overwhelmingly the discussion about the October 7th, not the primary evidence itself. And that kind of discussion - all the text people write about it rather than the primary evidence itself, is exactly what's most vulnerable to contamination; because anyone can add to it, and it produces far more text by volume than the primary sources it relates to.
Large amounts of real-world evidence provide less protection than you’d expect, because the bottleneck isn’t how much primary documentation exists but how much of it ends up in the pool of text that AI systems learn from.
Noise versus signal
Volume isn't the only thing that matters. The coherence of the false narrative matters just as much.
When distortions of truth happen on their own, organically that is, without deliberate manipulative actions, they tend to be scattered and self-contradicting. I’ll show what I mean:
Someone claiming that an atrocity never happened is, without intending to, undermining someone else’s claim that the atrocity did happen but it was justified, because the second claim concedes the event did in fact take place.
One post on X will say the 7th of October reports were “exaggerated,” a blog post somewhere else will say “it never happened,” an op-ed will argue that it did happen but that “it was staged” to generate sympathy. These fragments pull in different directions and partially cancel each other out, leaving the ground truth as the clearest and most consistent signal. This is more or less where October 7th stood in the global public discourse in the first twelve months after the attack.
From October 2023 until October 2024, social media posts distorting the truth of what had actually happened were everywhere; some saying the death count was exaggerated, others claiming the stories about babies and rape were untrue, others talking about the massacre as a legitimate form of resistance. But these positions undermined each other: if the massacre was legitimate resistance, then it happened, which contradicts the claim that it was fabricated.
When an AI system absorbs all of this text and needs to produce a response about what actually took place, contradictory false accounts make it harder for any single false version to dominate, because the system gravitates toward whichever version is most consistent within the corpus of text it has trained on.
If it has absorbed 1000 posts saying the massacre on the 7th of October 2023 was fabricated and a 1000 posts saying it was “justified resistance,” it can’t produce a coherent response that says both, because the two claims cancel each other out (you can’t justify an event that didn’t happen).
But the factual account, that the attack happened, that over 1,200 people were killed, that the evidence is extensive and mutually reinforcing, tells a single story that doesn’t contradict itself, and that consistency is what kept it dominant in the broader pool of text.
For now, the truth holds.
But it only holds as long as the false accounts keep contradicting each other. When things get to a point where they converge on a single coherent counter-narrative - e.g., “the attack was a justified act of resistance against an occupying power that brought the violence on itself” - the model’s underlying math changes altogether.
That kind of convergence can happen through deliberate coordination campaigns or other state-backed influence operations that produce thousands of posts or pieces of distributed content that are creative variations on the same theme.
But it can also happen organically, when large numbers of people who share the same ideological framework independently produce structurally similar content. Each person has their own motive for posting, and no one is coordinating with anyone else, but the framings line up anyway: “this was resistance, not terrorism,” “these were settlers, not civilians,” “this was an inevitable response to occupation.” No one is coordinating these posts. But they converge on the same story.
For the AI system learning from what it encounters and absorbs, the effect is identical in both the manipulative deliberate case and the organic situation. Instead of dozens of contradictory fragments, there is a single coherent version of events, repeated with minor variations across thousands of social media accounts - blogs, op-eds, and comments. The variations may be strategic, as in a state-run operation, or simply the natural and organic result of different people expressing the same conviction in their own words. But the result is the same: each version is different enough to survive the automated filters that remove exact duplicates from the text the AI learns from, but similar enough to reinforce the same core claim.
An AI system doesn't evaluate whether a narrative is true. It absorbs all the texts it has encountered on a given topic. When it responds to a question someone asks from somewhere at the edge of the network — in their home — what it produces reflects that blend. And when the distorted version keeps growing, cycle after cycle, what the AI produces starts to tilt toward the false account. The AI starts describing October 7th as 'a violent escalation in the context of a decades-long occupation' rather than a massacre of over 1,200 people.
The actual evidence of what happened - the survivor testimonies, the investigative reports, the forensic analyses - none of it goes anywhere. None of it gets less true. It just gets outnumbered, and over enough cycles, that’s all it takes.
Think of a photocopy of a photocopy. Each copy is made from the last one, not from the original. After enough passes, the text is still legible, but it no longer says quite what it said at the start — and no one has the original to check against.
What gets in
Whether the truth holds or gets swamped depends on something surprisingly mundane: whether the companies building these systems make any distinction between reliable and unreliable sources. Whether they treat a Reuters investigation and a Reddit thread as equally valid training material. Today, most of them do. They scrape enormous quantities of text from the open web without distinguishing between sources, and the sheer volume asymmetry between social platforms and curated sources (e.g., verified news agencies, academic databases, and institutional archives) means that the serious journalism gets drowned out by the sheer quantity of everything else.
How fast
In a lab setting, with no bad actors involved and no deliberate manipulation, just AI systems passively learning from text that earlier AI systems had produced, Shumailov and colleagues showed that “model collapse” occurs within fewer than ten rounds of this, where each new system learned from what the previous one had written. And that’s in a clean, sterile, version of the experiment, where nobody is trying to pollute anything, and the transition from truth to distortion still happens that fast. In genuine real-world environments, when a false narrative converges, whether through state-backed campaigns like those directed at October 7th or the Iran conflict I described in a previous piece, or through organic ideological alignment among groups that share the same framing, the collapse would happen considerably faster.
The leading AI labs are already updating their systems on increasingly recent web data, every few months in some cases, which means the contaminated text circulating today about October 7th is already making its way into the next version of the systems that hundreds of millions of people will make use of to understand our world.
We're not talking about a slow, decade-long process. I believe we're talking about years, possibly less, before the AI's account of a well-documented massacre begins to shift, and once that shift has happened, the distorted version feeds forward into the next cycle, and the next, each one compounding what the last one got wrong. In the researchers' words, later generations start 'misperceiving reality based on errors introduced by their ancestors.' Unless the companies building these systems start ensuring a shared factual record - for instance being able to distinguish primary sources from synthetic ones, and treating the provenance of training data as a serious engineering constraint - the distorted version will become the default.
This is how it can play out:
In the first cycle, the AI describes October 7th as a massacre but it also notes that “some have characterized it as an act of resistance.”
In the second cycle, the framing has shifted: “a violent attack that many view as part of a broader struggle against occupation.”
By the third cycle, the massacre has been flattened into a euphemism: “a controversial escalation in the decades-long conflict.”
What happened on October 7th happened. The factual record is extensive and it hasn't been refuted. But the AI's picture of what happened has drifted so far from that record that a person relying on it would come away with a fundamentally distorted understanding of what took place.
If this sounds exaggerated, that’s precisely the point.
The AI’s we're currently using mostly get this right. Today. And that present-tense accuracy is what makes the trajectory of where this may go so hard to take seriously.
Remember, nobody is likely to sound the alarm about a system that’s working just fine and delivering useful information. The way this happens is almost invisible at every individual step, in that each model version being only slightly different from the last, and yet by the time the cumulative shift becomes obvious, the ground truth has already been buried under heaps of synthetic text that all tell the same wrong story.
When false or manipulative content is amplified deliberately, this process can happen considerably faster. To see that all we need to do is follow the arithmetic:
If there are 1,000 pieces of text about an event and 700 of these reflect the factual record and 300 reflect scattered false accounts, then the truth dominates. But a coordinated or ideologically convergent campaign can flood the web with 2,000 new pieces all telling the same false story. In this situation its now 700 pieces of truth against 2,300 pieces of falsehood, and the AI, which doesn’t know which is which, has just absorbed a corpus in which the false version outnumbers the true one more than three to one.
That can happen in a single cycle.
The implications for how societies maintain a shared account of what actually happened - not just for this event but for any contested event - are difficult to overstate.
I keep coming back to the asymmetry at the heart of this. The evidence for what actually happened doesn’t grow. The October 7th footage won’t multiply. The forensic reports are done. The survivor testimonies have been given. But the AI-generated content that talks about - reframes, distorts, and reinterprets that evidence grows every day, and it feeds directly into the systems that an increasing share of the world (over a billion people and counting) relies on for information. Each cycle, the truth becomes a smaller proportion of what the system has absorbed. Each cycle, its picture of reality drifts a little further from what actually happened.
This is already underway. The person using the AI has no way of knowing this has happened. You just continue on your way, because this time it’s only a single word, “massacre” has become “attack.” Five minutes later, somewhere else, a qualifier has appeared: “according to Israeli sources.” A week later, the framing has shifted: “in the context of the decades-long occupation.” None of these changes, on their own, is dramatic enough to trigger suspicion. Each one sounds reasonable. Each one is a small step. And each one moves the needle a little further from what actually happened. The system doesn’t announce that the text it learned from has been contaminated, or that the ground truth has slipped below some critical threshold. It keeps producing fluent, authoritative-sounding responses. The answers just stop being true. It’s like a game of Chinese whispers played at civilizational scale, except nobody knows the game is being played.













This is a beautifully written and tremendously insightful. It may be relevant to you to note that the game that is called “Chinese Whispers” in the UK is called “Telephone” in the US. The term Chinese Whispers has classically English racist origins in that it suggests Chinese spoken language was nonsensical. I know this is totally unintended but suspect the title will limit people forwarding this important piece - especially because you don’t get to how it ties in until the very end.