Sources — The Marks It Leaves

Every figure in the essay, where it comes from, and how far it can be pushed.

How much

Europol, Facing reality? Law enforcement and the challenge of deepfakes, Innovation Lab, 28 April 2022. The “90% of online content synthetic by 2026” line is on page 4 and reads in full: “Experts estimate that as much as 90 % of online content may be synthetically generated by 2026.” Its endnote 4 cites one source — Nina Schick, Deepfakes: The Coming Infocalypse, Twelve/Hachette UK, 2020 — a trade book, not a study. Note that Europol’s own sentence says “online content”; the restriction to video belongs to the underlying claim in the book, not to Europol.

Kobak, González-Márquez, Horvát & Lause, “Delving into LLM-assisted writing in biomedical publications through excess vocabulary”, Science Advances 11(27), 2025. Peer-reviewed. DOI 10.1126/sciadv.adt3813; preprint at arXiv:2406.07016. 15,103,888 PubMed abstracts after cleaning. delves frequency ratio 28.0, underscores 13.8, showcasing 10.7. Lower bound of 13.5% for 2024 abstracts. ⚠️ The worst subcorpus — computational papers from China in Sensors — is 41% in the full text and rounded to 40% in the published abstract; the essay uses 40%. The authors’ limitation, quoted in the essay, is exact: “Our analysis is performed on the corpus level and cannot identify individual abstracts that may have been processed by an LLM.”

Sun et al., “Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media”, ACL 2025 Main Conference, Long Papers. Peer-reviewed. aclanthology.org/2025.acl-long.1120; preprint arXiv:2412.18148. Around 2.4M posts. Medium 1.77% → 37.03%, Quora 2.06% → 38.95%, Reddit 1.31% → 2.45%. ⚠️ These are end points of a time series ending October 2024, not corpus averages. The January 2022 readings function as a rough false-positive floor on the same populations, and they are low.

Pangram, “AI in Your Feed”, July 2026. Vendor study, not peer-reviewed. 1,002,627 posts across five platforms, collected through a browser extension in the feeds of users who opted in. LinkedIn long-form (250+ words) over 40% fully AI; X long articles 23.9% fully AI and 22.9% AI-assisted; Substack 21.9% AI-generated or AI-assisted; Reddit 4.4% overall, 11.6% of top-level posts, with replies 98.1% human. ⚠️ Pangram’s own chief executive describes the browsing figures as a lower bound, since people who install an AI-detection extension are not a random sample.

Originality.ai, LinkedIn studies, 2024–2026. Vendor studies, not peer-reviewed. The 81.2% headline counts posts of 100 words or more and classifies them as “Likely AI”. ⚠️ The same firm’s data lands on roughly 41% when the threshold is raised to 250 words and the question is narrowed to fully machine-written — which is why the essay treats the 41–81% range as a matter of definition rather than a disagreement between vendors.

Graphite, “AI now writes as many online articles as humans do”, May 2026. Not peer-reviewed, and methodologically the strongest of the commercial work: 55,400 randomly sampled Common Crawl URLs, English, article schema, 100+ words, January 2020 to March 2026, classified by three independent detectors — Pangram, Copyleaks and GPTZero — which agree rather than being averaged into one another. Published false-positive rates for all three below 2%, measured against a pre-ChatGPT corpus. Q1 2026: 49.9%. Note this revised an earlier single-detector estimate downward by 3.3 points.

Ahrefs, “What percentage of new content is AI-generated?”, April 2025. Vendor study using an in-house detector, not externally validated. ~900,000 English-language pages, one per domain. 74.2% contain AI content; 2.5% pure AI; 71.7% human–AI blend; 25.8% pure human. ⚠️ “Blend” is the detector’s classification, and the band runs from light AI editing to pages that are almost entirely machine. The essay says so.

Chen, Ye, Ferrara & Luceri, “Prevalence, Sharing Patterns, and Spreaders of Multimodal AI-Generated Content on X during the 2024 U.S. Presidential Election”, arXiv:2502.11248; ACM Hypertext ’25. Around 3% of text spreaders account for 80% of the AI-generated text shared. ⚠️ Two limits the essay carries: this measures sharing, not authorship, and the base is small — about 1.4% of texts in the sample were assessed as AI-generated. The peer-reviewed conference version is scoped to images, so the text figure is cited to the preprint.

Where the counting stops

DiResta & Goldstein, “How spammers and scammers leverage AI-generated images on Facebook for audience growth”, Harvard Kennedy School Misinformation Review, 15 August 2024, DOI 10.37016/mr-2020-151. Peer-reviewed. Preceded by a preprint, arXiv:2403.12838, 19 March 2024. The source for the unlabelled AI-generated image among the ten most-viewed Facebook posts of Q3 2023 (40 million views, 1.9 million reactions). It is not a prevalence study and must not be cited as one. 125 Pages, discovered by hand, included only if they posted more than fifty AI-generated images, identified manually. The authors’ own limitation is quoted in the essay: the Pages studied “are not necessarily reflective of how unlabeled AI-generated images are used on Facebook as a whole”. They further note the sample over-includes Pages that curated their output poorly, under-includes those that used generated images sparingly, and was overwhelmingly English-language.

⚠️ The most-viewed ranking originates in Meta’s own Widely Viewed Content Report for Q3 2023. The DiResta and Goldstein characterisation of it has been read at source; Meta’s underlying report has not, and is marked here as second-hand rather than verified. Anyone leaning on the top-ten framing should open the archived Meta release first.

Reis, de Freitas Melo, Garimella & Benevenuto, “Can WhatsApp Benefit from Debunked Fact-Checked Stories to Reduce Misinformation?”, arXiv:2006.02471v2, 3 June 2020. Preprint; a version appeared in Harvard Kennedy School Misinformation Review. Cited for the methodological constraint only, in the authors’ words: “Due to the private encrypted nature of the messages on WhatsApp, it is hard to track the dissemination of misinformation at scale.” The paper is about misinformation, not machine-generated text; the constraint it states applies to both, and the essay uses it for the constraint alone.

Meta, “Labeling AI-Generated Images on Facebook, Instagram and Threads”, 6 February 2024 and “Our Approach to Labeling AI-Generated Content and Manipulated Media”, 5 April 2024 (with in-page updates of 1 July and 12 September 2024), plus the current Misinformation Community Standard. Read at source. The basis for the statement that Meta’s labelling covers image, video and audio and does not name text. Note the February post’s own ceiling: “it’s not yet possible to identify all AI-generated content, and there are ways that people can strip out invisible markers.” Detection of other vendors’ marks is written throughout as intention — “we’re building”, “as they implement their plans” — not as a capability in service. The live Community Standard page is client-side rendered, so its current wording is recorded here as closely matching the verified 2024 text rather than as byte-verified.

Meta, Community Standards Enforcement Report, Q4 2025. Read at source for the absence: fourteen Facebook and twelve Instagram policy categories, none for AI-generated, synthetic or manipulated media. The EU DSA transparency filings were checked at index level only — the underlying PDFs were not opened, so “no standing metric” is verified for the enforcement report and provisional for the DSA filings.

Meta, “What We Saw on Our Platforms During 2024’s Global Elections”, 3 December 2024. Source of the 590,000 figure. It counts image-generation requests Meta’s own system refused, not content on the platforms — the essay says so because the distinction is the whole point of citing it.

Meta Q1 2025 results release, 30 April 2025. Source of “almost 1 billion monthly actives” for Meta AI. ⚠️ An executive quotation inside a press release, not a reported financial metric: no definition of “monthly active”, not in the tables, not audited, and not repeated in the four subsequent releases checked. Cited in the essay with that status attached.

⚠️ WhatsApp’s user total is not stated in this essay, deliberately. The widely repeated three-billion figure could not be traced to any Meta press release or newsroom post; the last figure verifiable on a Meta-owned channel is “more than two billion”, from blog.whatsapp.com on 12 February 2020. Rather than print a stale figure or an unsourced current one, the essay makes its argument without a user count.

Meta Q2 2026 results release, 29 July 2026. Source of the 3.60 billion figure: “DAP was 3.60 billion on average for June 2026”. Family Daily Active People is a reported company metric, not an estimate — but it counts people opening the apps, not messages sent and not words written. The essay uses it only for population and says so; no figure exists, from Meta or anyone else, for how much text those people produce or what share of it a machine drafted.

Meta Messenger end-to-end encryption by default, December 2023. Cited for the fact that Messenger sits in the same measurement position as WhatsApp. ⚠️ Recorded from the announcement date as widely reported; the primary Meta post was not opened in this session’s research, so treat the December 2023 date as unverified-at-source.

Regulation (EU) 2024/1689 (AI Act), Article 50, read from the consolidated text on EUR-Lex and corroborated against a second source. Article 50(2) binds providers and covers “audio, image, video or text”. Article 50(4) binds deployers, defines deep fakes as “image, audio or video content”, and reaches text only where it is “published with the purpose of informing the public on matters of public interest”, with an exemption where the content “has undergone a process of human review or editorial control”. Article 113 sets application from 2 August 2026; Article 50 sits in Chapter IV and is named in none of the earlier or later carve-outs.

Regulation (EU) 2022/2065 (DSA), Article 35(1)(k), read from EUR-Lex. Covers “a generated or manipulated image, audio or video”. Text is not included. Note this is a risk-mitigation measure under Article 35(1) — “reasonable, proportionate and effective” measures against risks identified under Article 34 — rather than an unconditional labelling mandate, and the essay describes it as one of those measures rather than as a labelling mandate.

Meta, “Introducing Facebook Verified”, 24 July 2026. Read at source. Both quotations in the essay are verbatim.

⚠️ One question is open rather than answered. Whether an AI-generated item has ever appeared in a past Widely Viewed Content Report top-twenty table could not be established: earlier quarterly editions are distributed only inside a signed, expiring archive that could not be opened here. The essay makes no claim either way. The DiResta and Goldstein finding about Q3 2023 stands on their paper, not on our reading of Meta’s release.

The absence of a Facebook or Instagram prevalence figure is a negative result, and a qualified one. Searches of the preprint literature for measurements of AI-generated text or images on Meta platforms returned nothing, while returning prevalence studies for Reddit, Medium, Quora, the open web and the scholarly record. That sweep used a single method and could not reach institutional grey literature published off the preprint servers — the research centres the two authors above are associated with are the likeliest place for a figure this search would have missed. The essay states the weaker claim accordingly: none found, not none existing.

Watermarking and provenance

Dathathri et al., “Scalable watermarking for identifying large language model outputs” (SynthID-Text), Nature 634, 818–823, 23 October 2024. Peer-reviewed. Tournament sampling over candidate tokens, seeded by a secret key and the preceding n-gram.

Anthropic, “How Claude’s text watermarking works”, 14 August 2026. Company statement. Identifies the technique as a version of SynthID-Text, applied worldwide at launch. The quotation in the essay is verbatim: “Nothing is added to the text and there are no hidden characters.” A detection interface is described as forthcoming and is not published. ⚠️ Press coverage dates the announcement to 11 August; 14 August is the mechanics explainer.

OpenAI text watermarking. ⚠️ The 99.9% accuracy figure and the finding that nearly 30% of surveyed users would use ChatGPT less come from internal documents reported by the Wall Street Journal, 4 August 2024 — not from a published measurement. OpenAI’s own statement claims only that the method is “highly accurate”. Nothing is deployed for text.

EU AI Act, Regulation (EU) 2024/1689. Article 50(2) requires machine-generated content to be “marked in a machine-readable format and detectable as artificially generated or manipulated”, with technical solutions that are “effective, interoperable, robust and reliable as far as this is technically feasible”. Applies from 2 August 2026 under Article 113. ⚠️ Penalties for breaching Article 50 are set by Article 99(4)(g) at up to €15,000,000 or 3% of worldwide annual turnover, whichever is higher. The widely-quoted €35,000,000 figure belongs to Article 99(3) and pairs with 7%, applying to the practices Article 5 prohibits outright. The Commission’s Code of Practice on Transparency of AI-generated Content was finalised in June 2026 and asks providers to give third parties detection access; it is voluntary, while the Article 50 obligations are not.

C2PA specification 2.3 (5 January 2026) and 2.4. The unstructured-text embedding transcodes manifest bytes onto the 256 Unicode variation selector code points. ⚠️ L2/26-042, “Embedded Metadata in ‘Plain’ Text”, Peter Constable and Joshua Hadley, 13 January 2026 — a paper submitted to the Unicode Technical Committee by named authors, not a position adopted by the Unicode Consortium — argues the scheme uses valid characters in a non-conformant way under §23.4 of the Unicode Standard. The committee opened a liaison with C2PA and revised its own conformance language. The essay describes this as a working disagreement between standards bodies, which is what it is.

Markers

Reinhart et al., “Do LLMs write like humans? Variation in grammatical and rhetorical styles”, PNAS 122, 2025. Peer-reviewed. Finds systematic differences persisting across model families and sizes; instruction-tuned models show a noun-heavy, informationally dense style.

Herbold et al., “A large-scale comparison of human-written versus ChatGPT-generated essays”, Scientific Reports 13, 18617, 2023. Peer-reviewed. Fewer discourse and epistemic markers, more nominalisations. ⚠️ On lexical diversity this paper reports different directions for different model versions — an earlier ChatGPT below human writers, a later one above. Cite the version, not the paper. ⚠️ The essay originally claimed these two studies had “replicated” a finding about fewer hedges. They had not. Neither paper measures hedging as such; one hedging-adjacent measure in the PNAS work runs the other way for one model. The two overlap on nominalisation and density. The essay was corrected.

Czuma, “Em-ergence of the em-dash”, arXiv:2606.29540, 28 June 2026. Preprint — not peer-reviewed. 69,632 medRxiv clinical preprints. ⚠️ The endpoint is presence of at least one em-dash in the Discussion section, not a rate across the paper, and the 4.23% → 11.58% comparison is a pre/post-ChatGPT split, not calendar years. 2025 alone: 20.3%. The paper’s own limitation, quoted in the essay: it “is a population-level indicator, not a per-paper detector of LLM use.”

Freeburg, “The Last Fingerprint: How Markdown Training Shapes LLM Prose”, arXiv:2603.27006, March 2026. Preprint — not peer-reviewed. Twelve models. GPT-4.1 10.62 em-dashes per 1,000 words; Claude Opus 4.6 9.09; Gemini 2.5 Pro 3.53; GPT-5.4 1.43; Llama 3.1 0.00. Human baseline 3.23, range 0.33–17.12. ⚠️ The human baseline rests on eight published essays, 57,232 words — not a population. The suppression figure (Claude 9.09 → 0.19) follows a markdown ban, not an em-dash ban; an explicit em-dash prohibition failed to shift some models.

Liang, Yuksekgonul, Mao, Wu & Zou, “GPT detectors are biased against non-native English writers”, Patterns 4(7), 2023; preprint arXiv:2304.02819. Peer-reviewed. Seven detectors, 91 TOEFL essays, 88 US eighth-grade essays.

⚠️ Two of the four figures the essay uses appear only in the preprint. The published article prints 61.3% and 11.6%, and describes the eighth-grade results qualitatively with the values in Figure 1. The preprint prints 61.22%, 5.19%, 11.77% and 56.65%. The essay’s 5% and 57% are cited to the preprint. The TOEFL essays were collected from an educational forum rather than gathered under observation, and 91 is a small base for a finding this load-bearing — both stated in the essay.

Jakesch, Hancock & Naaman, “Human heuristics for AI-generated language are flawed”, PNAS 120(11), 2023. Peer-reviewed. 4,600 participants, 7,600 self-presentations, 50–52% accuracy. The heuristics finding is verbatim: participants associated “first-person pronouns, use of contractions, or family topics with human-written language”.

GPTZero support documentation. The company states that as of autumn 2023 it no longer uses perplexity and burstiness for detection, having moved to a deep-learning architecture, and that they remain one of seven indicators. ⚠️ Burstiness itself is not a vendor invention — it is established statistical language processing (Church & Gale 1995; Katz 1996; Madsen, Kauchak & Elkan, ICML 2005). What the vendor contributed is a particular operationalisation.

Fredrick & Craven, Frontiers in Education, 2025. Lexical diversity, ChatGPT versus second-language student essays: TTR 0.69 vs 0.61, MTLD 118 vs 66.56, Voc-D 105 vs 69.73. ⚠️ Three measures, fifty essays a side, one programme, no inferential statistics reported. Direction is model-dependent — Herbold et al. found an earlier ChatGPT version lower than human writers on diversity, which is why the essay makes the model-dependence the point rather than the direction.

Yakura et al., arXiv:2409.01754. Preprint. 737,083 hours of unscripted podcast speech across 824,634 episodes; AI-associated vocabulary rising in spontaneous human speech after ChatGPT.

Ramez Naam, writing on his own site (rameznaam.com), argues that AI will be plural and multi-polar rather than a single dominant system, that it is being democratised faster than earlier technologies, and that this plurality favours human freedom. The essay summarises rather than quotes him, and the summary is of the position rather than of any single sentence. ⚠️ The observation that the frontier changed hands fourteen times between four companies in twelve months is his, repeated as his and not verified here — the companion piece says so where it uses it.

Defeating the marks

Krishna, Song, Karpinska, Wieting & Iyyer, “Paraphrasing evades detectors of AI-generated text” (DIPPER), NeurIPS 2023; arXiv:2303.13408. Peer-reviewed. DetectGPT accuracy 70.3% → 4.6% at a fixed 1% false-positive rate. ⚠️ That is the most aggressive single paraphrase setting; milder passes give 28.7%, 15.4% and 8.7%. The paper’s retrieval defence detects 80–97% of paraphrased text against a database of 15M generations.

Sadasivan, Kumar, Balasubramanian, Wang & Feizi, “Can AI-Generated Text be Reliably Detected?”, arXiv:2303.11156; TMLR 2025. ⚠️ Cite version 4. The watermark figures the essay uses — 99.8% → 80.7% after one DIPPER paraphrase on a watermarked OPT-13B, and 99.3% → 9.7% under recursive paraphrasing — are absent from versions 1 and 2. These figures are Sadasivan’s, not Krishna’s; the DIPPER paper’s own watermark result is 100% → 57.2%.

⚠️ The theoretical bound is AUROC ≤ ½ + TV(M,H) − TV(M,H)²/2, and the authors explicitly disclaim the strong reading. Their words: “Our impossibility result does not imply that detection performance will necessarily become as bad as random, but that reliable detection may be unachievable.” Version 4 replaced “impossibility” with “hardness”. The essay states it as the authors do.

He et al., “Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language Models”, ACL 2024; arXiv:2402.14007. Peer-reviewed. True-positive rate at fixed 0.1 false-positive rate, English–Chinese–English round trip: SIR 0.940 → 0.825, KGW 0.992 → 0.776, Unbiased Watermark 0.913 → 0.263. Generating in the pivot language and translating afterwards: 0.230, 0.213, 0.166 — the “around a fifth” in the essay is the paper’s own result, not an extrapolation.

Dugan et al., “RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors”, ACL 2024. Peer-reviewed. Over 6 million generations, 11 models, 8 domains, 11 attacks. At a fixed 5% false-positive rate the homoglyph attack costs Binoculars 41.9 points and Originality.ai 75.7; synonym substitution costs Binoculars 36.1. ⚠️ Attribute each drop to its own detector — several improve under paraphrase.

What was checked and how

Every figure above was checked against its primary source — publisher full text, ACL Anthology, arXiv PDFs, official EU documents and company statements, rather than secondary write-ups.

That checking was done by AI agents, working separately from the one that drafted the essay, and it should be described as what it was rather than dressed up as independent human review. Several ran the same questions by different methods. Where they disagreed, the disagreement was settled by which one produced a document rather than by which sounded more certain — twice, an agent reported it could not find a study that another had already located with a verbatim quotation and a working identifier.

The pass found ten errors in the draft. A penalty figure attributed to the wrong tier of the EU AI Act. A claim that the Unicode Consortium had formally objected to something, when the objection was an individual paper put before its technical committee. A marker described as “replicated” across two studies that had not replicated it. A statistic about round-trip translation with no source behind it anywhere. All were corrected before publication, and several are recorded above so a reader can see what the essay used to say.

The essay argues that a declaration costs nothing and settles what no test can. It would be a poor advertisement for that argument if this page implied a person had done work a machine did. Two negative findings — that no peer-reviewed measurement establishes a majority of long-form social media as AI-written, and that no round-trip-translation figure of the kind widely quoted exists — were each checked twice by different methods before being relied on.

Alongside: back to the essay