Every study the essay rests on, what it actually found, and where it stops.
The research this essay was built from arrived carrying eight
placeholder citation markers of the form [web:31],
[web:66] — search-result handles that resolve only inside
the session that produced them. They are not references: no author, no
title, no venue, no year, no URL, and nothing a reader could follow.
Every one was traced to its primary source and read there. All eight resolve to real, checkable work — six peer-reviewed, one an institutional report, one a preprint. Nothing was invented.
That checking was done by AI agents working separately from the one that drafted the essay, and it should be described as what it was rather than dressed up as independent human review. It changed the argument in three places, listed at the foot of this page.
Vishakh Padmakumar, He He, “Does Writing with Language Models Reduce Content Diversity?”, ICLR 2024. arXiv:2309.05196. Peer-reviewed.
⚠️ The boundary matters and the essay states it. The base model produced no significant effect. This is a property of feedback-tuned assistants, not of language models as such. Also: paid experienced writers, argumentative essays only, one interaction paradigm, two 2022-vintage models. Effect sizes are small in absolute terms.
Dhruv Agarwal, Mor Naaman, Aditya Vashistha, “AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances”, CHI 2025. arXiv:2409.11360. Peer-reviewed.
⚠️ The authors note their Indian participants were crowdworkers, likely more familiar with American norms than the general population — so the real cultural distance may be larger than measured. The Indian sample was also 71.7% male against a roughly balanced US sample.
Elliot Kim, Avi Garg, Kenny Peng, Nikhil Garg, “Correlated Errors in Large Language Models”, ICML 2025. PMLR 267:30038–30066. arXiv:2506.07962. Peer-reviewed.
🔴 The baseline is the part that gets dropped, and the paper states it plainly. Chance agreement is 33% on HELM (three wrong options) and 12.7% on HuggingFace. So the 60% figure is 1.8× chance and the 42.3% figure is 3.3× — the smaller number is the stronger finding. The essay gives both with their baselines; almost everything quoting this study gives neither.
Two findings that matter more than the headline, both from the paper’s own abstract:
⚠️ The measurement is multiple choice with ground-truth labels — four options on HELM, more on the HuggingFace set — not open-ended generation, which is what someone consulting several assistants actually does. The authors say richer evaluation remains future work. This limit is stated in the essay because the study cannot carry the weight some readings put on it.
Stanford Institute for Human-Centered AI, The 2026 AI Index Report, Research and Development chapter. Not peer-reviewed — an institutional annual report, reporting 2025 data.
Verified two ways that do not share a method: the AI Index chapter itself, and independent reporting of the same figures by IEEE Spectrum. The compute figures originate with Epoch AI, whose data the Index reports.
Zhivar Sourati and colleagues, “The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models”, arXiv:2502.11266.
🔴 PREPRINT — not peer-reviewed. No venue, single version, unpublished. The essay does not lean on it and labels it.
⚠️ The authors concede raised Type I error risk from exploring multiple lags in the Granger-causality tests underpinning the observational study, which is the one that would carry any population-scale inference. Predictive power was reduced, not eliminated — and the 6% drop is from an already low base.
Sarah Fitterer, Dominik Gangl, Jannes Ulbrich, “Testing English News Articles for Lexical Homogenization Due to Widespread Use of Large Language Models”, ACL 2025 Student Research Workshop. DOI 10.18653/v1/2025.acl-srw.95. Peer-reviewed — but note the track: a seven-page student workshop paper, not a main-conference long paper.
Roughly 30,000 English news articles from 2018 against 30,000 from 2024, drawn from the News on the Web corpus.
| Measure | 2018 | 2024 | Direction |
|---|---|---|---|
| MATTR | .88011 | .88121 | flat |
| Maas | .01469 | .01482 | flat |
| MTLD | 214.45 | 254.65 | up |
| LLM-style word ratio | 0.230% | 0.347% | up, significantly |
🔴 Three things about this study, all of which the essay’s earlier draft got wrong or omitted.
That third point is why this does not settle the question in either direction. It is a measurement failure reported as one.
⚠️ Further limits: no significance tests, only a split-sample variation baseline; the LLM-style word list is borrowed from a study of biomedical abstracts, a different register; news only; and news outlets have editorial AI policies, so AI content in the 2024 sample may be sparse by design.
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, Yarin Gal, “AI models collapse when trained on recursively generated data”, Nature 631, 755–759 (2024). DOI 10.1038/s41586-024-07566-y. Peer-reviewed.
Recursive training on generated data causes “tails of the original content distribution [to] disappear” — what is rare drops out across generations.
⚠️ The contested assumption, which the essay names. The experiments replace the training data each generation. Gerstgrasser and colleagues showed that when real data accumulate alongside synthetic data — closer to how the web actually works — test loss does not diverge. There is also a direct critical note at arXiv:2410.12954. The original experiments used small models (OPT-125m) and fine-tuning rather than training from scratch. Retaining 10% of original data across ten generations produced only minor degradation.
The essay’s conclusion — a reason for vigilance rather than a prediction of collapse — survives the critique, which is why the critique is stated rather than omitted.
Badr AlKhamissi, Muhammad ElNokrashy, Mai Alkhamissi, Mona Diab, “Investigating Cultural Alignment of Large Language Models”, ACL 2024 (long paper). DOI 10.18653/v1/2024.acl-long.671. Peer-reviewed.
Alignment improves when a model is prompted in the relevant dominant language and pretrained on a more culturally appropriate language mixture. Measured by simulating sociological surveys across Egypt and the United States, in Arabic and English.
⚠️ This cuts both ways and the essay says so. Misalignment becomes more pronounced for underrepresented personas and on culturally sensitive topics — precisely where the concern is sharpest. And the authors’ own recommendation is a data-layer intervention, which is the essay’s argument rather than a rebuttal of it. Note also that “cultural alignment” here means agreement with survey response distributions, which says nothing about register or provenance.
These are the claims with real exposure, and an earlier draft carried neither a citation nor an accurate quotation. Both are corrected here.
OpenAI submission to the White House OSTP, 13 March 2025 — response to the Request for Information on the AI Action Plan, addressed to the National Coordination Office. Eight pages. It contains a section headed Advancing democratic AI (p.3) and one headed Export Controls: Exporting Democratic AI, proposes a Tier I/II/III country structure, and uses both “democratic rails” and “American rails” (p.6).
⚠️ Why this document and not another. It is cited because it is the clearest written instance of a general move, not because the move belongs to one country. The essay says so where the quotation appears, and notes the direction deployment actually runs: the open-weight models most organisations adopt — often specifically to avoid depending on a US vendor — are increasingly Chinese, and they fail the essay’s fourth condition for identical structural reasons. No comparable published statement from another bloc was located in this research pass, which is a limit of the search rather than evidence that none exists.
🔴 An earlier draft of this essay quoted a sentence from it incorrectly. It rendered “American-led AI built on democratic principles continues to prevail over CCP-built autocratic, authoritarian AI” as a flat assertion. The original is conditional and begins earlier: the AI Action Plan “can ensure that” American-led AI continues to prevail. That is a policy aim, not a claim of fact, and the clipped version came from the covering paragraph rather than from the section it was attributed to. Quoting a claim without the words that qualify it is precisely what this essay criticises elsewhere, so the quotation has been removed entirely and replaced with the document own section headings, which need no interpretation.
Koster, Balaguer, Tacchetti, Weinstein, Zhu, Hauser, Williams, Campbell-Gillingham, Thacker, Botvinick & Summerfield, “Human-centred mechanism design with Democratic AI”, Nature Human Behaviour 6, 1398–1407 (2022). DOI 10.1038/s41562-022-01383-x. Peer-reviewed. Abstract, verbatim: a pipeline “in which reinforcement learning is used to design a social mechanism that humans prefer by majority”, which “successfully won the majority vote”. The essay characterises this as a majority-preference selector, which is what the paper says it is. Note it designs a redistribution mechanism, not outputs — the essay does not claim otherwise.
Anthropic text watermarking. The declaration at the top of the essay states that Claude began watermarking its text output in August 2026. Source: Anthropic’s own published statement, “How Claude’s text watermarking works”, 14 August 2026, which identifies the technique as a version of SynthID-Text and says it is applied worldwide. ⚠️ This is a company statement about its own product, days old at the time of writing, with no independent verification available — a detection interface is announced and unpublished. It is stated here as what it is.
Ramez Naam. The plural-AI argument the essay engages is set out by Naam on his own site, arguing that AI will be plural and multi-polar rather than a single dominant system and that this favours human freedom. The essay summarises the position rather than quoting it, and does not attribute to him the uses the argument has been put to by others.
The essay’s section on compression removing provenance is not an original observation and does not present itself as one. The frameworks that articulated it, and articulate it more precisely:
⚠️ WAI 262 is not a data governance framework and the essay does not cite it as one. It is named because it is why these questions have a settled legal and constitutional context in Aotearoa rather than an emerging one. An earlier draft grouped it with TMR and CARE as though the three were the same kind of instrument; they are not.
⚠️ An earlier draft also said these frameworks “articulated it first”. That is an unsourced claim of historical precedence over a long literature on situated and tacit knowledge, and it has been removed. What can be said without a priority claim is that they state the point more precisely than this essay does, which is true and sufficient.
These are named because the essay makes a claim in their territory and should follow them rather than paraphrase them. The essay also states plainly that the failure they describe — extraction without consent — is not solved by plurality, and is therefore a different problem from the one the essay argues about.
Three things, recorded because an essay about provenance should show its own.
Three further corrections came from the pre-publication review: the 60% correlated-error figure had been attached to the wrong dataset (it is HELM’s 71 models, not the HuggingFace leaderboard’s 349); a figure’s alternative text contradicted the figure’s own labels; and a diagram listed shared hardware among the causes of correlated error, which nothing cited supports.
A further change was structural rather than factual: the argument’s most useful output — a procurement test for correlated failure — was buried in the fifth row of a seven-row table. It is now a section of its own, because it is the one thing here that somebody can act on without waiting for a state or a standards body.
Alongside: the essay · questions and answers · glossary · slides · The Marks It Leaves