How much of what you read was built with a machine, what that leaves behind, and why saying so settles what no test can.
A question worth asking before any of the others: how much of what you read in a day was written by a person?
Not a rhetorical question. It has an answer, or at least the beginnings of one, and the answer is more interesting than either of the things people usually say about it.
The figure that circulates most widely is that ninety per cent of online content will be synthetic by 2026. It appears in Europol’s 2022 report on deepfakes, and from there in a great many articles. Follow its footnote and it leads to a single 2020 trade paperback. No measurement was ever performed. The word in Europol’s own sentence is estimate. And the sentence in the book it rests on was about video — “90 percent of the video content online is going to be synthetic” — so a claim about pictures has been laundered into a statistic about prose.
So begin somewhere else. Here is what has actually been counted — and, in the last four rows, what has not.
| Where | Share of long-form posts assessed as AI-written | Who counted, and how |
|---|---|---|
| 41% entirely machine-written; 81% with AI involved somewhere | Two vendors, two thresholds — the range is definitional, not a disagreement | |
| X (long articles) | 23.9% fully AI; 22.9% AI-assisted | Pangram, ~1M posts in opt-in users’ feeds |
| Quora | ~39% | Sun et al., ACL 2025, peer-reviewed, 2.4M posts |
| Medium | ~37% | as above |
| Substack | ~22% with AI involved or AI-assisted | Pangram |
| 4.4% overall, 11.6% of top-level posts, 2.5% peer-reviewed | Vendor and peer-reviewed, and note replies are 98% human | |
| New web articles | ~50% mostly machine-written | Graphite, 55,400 random URLs, 100+ words, three detectors that agree |
| Biomedical abstracts | at least 13.5% in 2024, rising | Kobak et al., Science Advances, no detector used at all |
| not counted | No published measurement found. The nearest study covers 125 hand-picked Pages and is not representative, as its authors say | |
| not counted | No published measurement found. The Facebook study does not cover it | |
| cannot be counted from outside | End-to-end encrypted. No third party can survey it, by design | |
| Messenger | cannot be counted from outside | End-to-end encrypted by default since 2023 |
The four blanks are Meta’s apps, which the company reports 3.60 billion people opening daily. They are blank because of how those platforms are built, not because nobody looked — which changes what the rest of the table means, and is taken up in Where the counting stops.
One caveat covers most of the table before any single row does. Three of these rows come from one detection vendor, measuring posts in the feeds of people who had installed its browser extension — and people who install an AI-detection extension are not a random sample of anybody. The firm’s own chief executive calls the resulting figures a lower bound. Read the whole table as indicative rather than as a census.
Two further things matter more than any single row.
The first is what the range in that top row is made of. One firm puts LinkedIn long-form at 81%, another at just over 40%, and the obvious reading — two companies contradicting each other — is wrong. They are measuring different things. The high figure counts posts of 100 words or more and asks whether AI was likely involved; the low one counts posts of 250 words or more and asks whether the post is entirely machine-written. Apply the stricter threshold to the first firm’s own data and it lands on 41%, alongside the second.
Which is worth more than a disagreement would have been. The headline numbers in circulation are not comparable with each other, and almost nobody quoting them says which question was asked. Align the definitions and the estimates converge — so the figure to carry is roughly two in five long-form LinkedIn posts written entirely by a machine, with a larger share again involving one somewhere.
The second is that the highest well-founded figures are not about social media at all. They are about newly published web articles, where roughly half now appear to be machine-written, and about academic abstracts, where the best-founded measurement in the field lives.
That measurement is worth pausing on, because of how it was done. Kobak and colleagues counted words across 15.1 million PubMed abstracts, using a method borrowed from the study of excess mortality: work out how often a word should appear in 2024 based on how often it appeared before, then measure the gap. The word delves turned up 28 times more often than expected. Underscores, nearly 14 times. No detector was involved. Nothing was classified. They simply counted, and the counting says at least 13.5% of 2024 abstracts had been through a model — a floor, not a ceiling, and around 40% in the worst subcorpus they looked at.
So: common, unevenly, and more common in some places than the loudest claims suggest. There is also a shape to it that gets lost in the averages. In a study of political posts on X during the 2024 US election, about 3% of the accounts that shared machine-written text accounted for 80% of what was shared. Two things about that are worth keeping straight: it measures sharing rather than writing, and the base was small — only around 1.4% of election-related posts contained machine-written text at all. What it shows is not how much there was. It is that whatever there was came overwhelmingly from very few accounts.
And one more decomposition, which turns out to be the whole argument in miniature. A widely-quoted finding that 74% of new web pages are AI-generated breaks down, in its own data, into 2.5% pure machine, 25.8% pure human, and 71.7% some mixture of the two. The headline counts any trace of AI at all.
That middle bucket is where the interesting question lives, and it is worth being careful about how wide it is: it runs from pages with a light pass of AI editing to pages that are 99% machine with a human hand at the end. Calling all of it collaboration would flatter a good deal of it. What it does establish is that a binary question has almost no purchase — the overwhelming majority of pages are neither one thing nor the other, and no test that returns human or machine can describe them. A sentence from whoever made the page could.
Go back up to that table, because every row in it has one thing in common. LinkedIn, X, Quora, Medium, Substack, Reddit, the open web, PubMed. An outsider can read all of them. That is not a coincidence in the subject matter; it is the selection rule. The table is a map of what can be crawled, and it has been quietly standing in for a map of where people write.
The largest surfaces are missing. There is no row for Facebook, none for Instagram, none for WhatsApp, none for Messenger.
How much writing sits behind that absence is itself uncounted, and it is worth being exact about which part is known. Meta reports 3.60 billion daily active people across its apps for June 2026. That is a figure for people, published by the company. It is not a figure for how much they write, still less for how much of it a machine drafted — nobody publishes either. So the shape of the claim is: an enormous population, an unknown volume, and no measurement of the share.
They are missing because no one has published a count — none this essay’s search could find — and in one case because no outside party could produce one.
The search behind that sentence, and who ran it, are described on the sources page. Searches of the preprint literature turn up prevalence studies for Reddit, Medium, Quora, the open web and the scholarly record, and none for either Meta platform, in text or in images. That is a negative result from one method rather than a proof of absence, and it should be read as the weaker claim.
The nearest thing to a study is this essay’s own argument in miniature.
In 2024 Renée DiResta and Josh Goldstein published an investigation into AI-generated images on Facebook — a preprint in March, then peer-reviewed in the Harvard Kennedy School Misinformation Review that August. They found spam and scam Pages using generated images at scale to build audiences, and one such image, unlabelled, among the ten most-viewed posts on Facebook in the third quarter of 2023, with 40 million views and 1.9 million reactions.
That is a real finding and a startling one. It is also not a prevalence figure, and the authors are unusually direct about it. They studied 125 Pages, found by hand, included only those posting more than fifty generated images, and identified the images manually. Their own sentence:
The Pages we studied are not necessarily reflective of how unlabeled AI-generated images are used on Facebook as a whole.
So a qualitative investigation of 125 hand-picked Pages cannot tell you what share of Facebook is machine-made, and its authors say so in the paper. If you meet a percentage for Facebook attributed to this work, it has been manufactured somewhere between the paper and you — the same route the ninety-per-cent claim travelled at the start of this essay, and a reminder that the laundering is not a historical curiosity.
WhatsApp is a different case, and a more interesting one. It is not unmeasured. It is unmeasurable, and by design.
The messages are end-to-end encrypted, which is the property most people would want it to have. It also means nobody outside the conversation can survey what is flowing through it — no researcher, no detector, no regulator, no third party. Meta itself could only measure by inspecting text on the handset before it is sealed, which is the thing the encryption exists to prevent. So a WhatsApp figure has either not been produced, or was produced by reading people’s messages on their own phones. Researchers on the platform describe the same wall from the other side. Here in a 2020 study of misinformation, meeting the constraint for a different subject:
Due to the private encrypted nature of the messages on WhatsApp, it is hard to track the dissemination of misinformation at scale.
Messenger has been end-to-end encrypted by default since December 2023 and sits in the same position. What can be studied on either is public groups a researcher has joined, which is a self-selected sample discovered by invitation and cannot support a claim about the platform. Anyone quoting a percentage of WhatsApp traffic as machine-written is quoting a number that could not have been produced.
There is a second reason, and it is not about researchers.
Meta has been labelling AI content since February 2024 — the “AI info” badge, applied when it identifies industry-standard indicators or when a user declares it. Read every scope-defining sentence the company has published, though, and the same three words recur: image, video, audio. The February announcement, the April policy, the July and September revisions, the Community Standard live now. In none of them is generated text named as covered.
So on the platforms where most people do their writing, the platform’s own provenance system does not extend to writing.
Nor is there a number. Meta’s enforcement reporting runs to twenty-six policy categories across Facebook and Instagram — nudity, harassment, fake accounts, spam, violence — and not one of them is synthetic or AI-generated media. The nearest thing the company has published to a quantity is 590,000: the number of requests it refused to generate images of named politicians around the 2024 elections. That counts its own generator declining prompts. It says nothing about what is circulating.
Set that against scale. Meta’s assistant sits inside Facebook, Instagram, WhatsApp and Messenger, and Mark Zuckerberg told investors in April 2025 it had “almost 1 billion monthly actives” — a remark in a results release rather than an audited metric, and it has not been repeated since. There is no published figure, anywhere, for how much of what that assistant drafts is subsequently posted or sent.
In July 2026 Meta introduced Facebook Verified. Its stated reasoning:
As AI makes it easier to generate content, profiles, and messages, we want to provide a way for you to know there is a real person on the other side of a profile.
The badge “means a profile belongs to a real person — someone who completed selfie verification”.
The categories that sentence names are profiles and messages — the two the labelling policy has never covered. The image labelling continues; alongside marking the machine, Meta has begun verifying the person.
Identity verification buys assurance by spending anonymity, and there are people for whom anonymity is not a convenience. A declaration costs the writer nothing and reveals nothing. A selfie check costs a name.
Three things follow.
Back on the open internet, where the counting happens, machine-written text does leave marks. They fall into four kinds, and the kinds are not equally worth having. Three are properties of the writing. The fourth is not, and it is the one you can see.
The strongest signal is the one somebody meant to leave. This became real very recently.
Google’s SynthID has been running on Gemini since 2024, published in Nature, and it works by nudging which word gets chosen at each step — a secret key and the preceding few words decide a tournament between candidate words, and the winners are biased in a way that shows up statistically across a long enough passage. Anthropic announced the same family of technique for Claude in mid-August 2026, publishing the mechanics on the 14th, and is applying it worldwide rather than only in the jurisdiction that required it. OpenAI’s case is the instructive one. According to internal documents reported by the Wall Street Journal, it had a text watermarking tool ready for about a year and measured at 99.9% accuracy, and did not ship it. Among the considerations reported was an internal survey in which nearly 30% of users said they would use ChatGPT less if it did. OpenAI’s own published statement claims only that the method is “highly accurate”, and the tool remains undeployed.
What forced the change was European law. The EU AI Act’s transparency obligations applied from 2 August 2026, with penalties for breaching them reaching €15 million or 3% of worldwide turnover, whichever is higher. (The €35 million figure that circulates alongside this belongs to a different tier of the Act, covering practices it prohibits outright.)
Now the part that decides what any of this is worth to a reader.
The Act requires the mark to be machine-readable. It does not require anyone but the vendor to be able to read it.
The nearest thing to a counter-argument is a single word. The Act asks that the technical solutions be “effective, interoperable, robust and reliable as far as this is technically feasible” — and interoperable sounds like it should mean a mark one party applies can be read by another. In practice nothing yet defines what that requires, the standards work is unfinished, and no vendor has published a verifier on the strength of it.
There is a Commission code of practice that goes further, asking providers to give researchers, journalists, authorities and civil society the means to check. It was finalised in June 2026 — and it is voluntary, which the Commission says in terms while noting that the Article 50 obligations themselves are not. So the binding part requires marking, and the part that would let anyone else read the mark is the part nobody has to sign.
Google’s detector exists and is a waitlist — for journalists and researchers, for Google’s own content only. Anthropic’s detection interface is announced and unpublished. OpenAI has nothing deployed to detect. So the marks are real, they are being applied at scale, and there is no way for you to check one. The mark exists; the key does not travel with it.
The statute divides along the same seam. Article 50 places duties on two different parties, and they are not the same duty.
On whoever built the model, Article 50(2) — providers of systems generating synthetic “audio, image, video or text” must ensure outputs are “marked in a machine-readable format and detectable as artificially generated or manipulated” — the obligation quoted above. It does not apply where a system performs “an assistive function for standard editing” or does not substantially alter the input.
On whoever runs the platform, Article 50(4) — deployers must disclose deep fakes, defined as “image, audio or video content”. Text appears only in a much narrower clause, covering text “published with the purpose of informing the public on matters of public interest”, and the provision does not apply where the content “has undergone a process of human review or editorial control” and a person “holds editorial responsibility for the publication”.
The Digital Services Act divides in the same place. Among the risk-mitigation measures very large platforms may take against identified systemic risks, Article 35(1)(k) names “a generated or manipulated image, audio or video”. Text is not in it.
Article 50(2) names text at the point of manufacture. Neither Article 50(4) nor the Digital Services Act provision names text at the point of carriage. The statute obliges the maker to mark, and obliges nobody downstream to look.
There is one scheme built the other way. C2PA Content Credentials are a signed record of provenance that anyone can verify with public tooling, no vendor secret needed — the difference matters and should not be blurred. But its text encoding was only published in January this year, has essentially no deployment by any major provider, and dies to a screenshot or to any tool that tidies up character encoding. (It does survive an ordinary copy and paste — the characters were chosen for that.) When present it proves something. When absent it proves nothing at all.
The way it works for text is to hide the record in invisible Unicode characters, and that has not gone uncontested. A background paper put to the Unicode Technical Committee argued the scheme uses valid characters in a way the Standard does not sanction, noting that earlier attempts to carry data in invisible formatting characters have been deprecated for breaking text processing. The committee’s own response was to open a liaison with the C2PA and to revise the relevant conformance language in its own Standard — a working disagreement between two standards bodies rather than a verdict.
Below the deliberate marks sit the accidental ones — the habits of machine writing that somebody has counted.
The excess-vocabulary work above is the best of these, and it comes with a caveat its authors state plainly: the method “cannot identify individual abstracts that may have been processed by an LLM.” It estimates a population. It cannot point at anyone.
The em-dash is the same story, and its scope is narrower than the way it gets quoted. Across 69,632 clinical preprints, the share whose discussion sections contained at least one em-dash went from 4.23% before ChatGPT’s release to 11.58% after it — and, taking 2025 on its own, to 20.3%. The paper states its own limit: this “is a population-level indicator, not a per-paper detector of LLM use.”
And when you look at machines individually rather than in aggregate, the em-dash stops being a property of AI at all. Measured across twelve models: GPT-4.1 uses 10.6 per thousand words, Claude Opus around 9, Gemini 2.5 Pro 3.5, GPT-5.4 1.4, and Llama 3.1 uses none whatsoever. The human baseline in the same study was 3.2 — sitting in the middle of that range, with the eight human texts they measured running from 0.3 to 17. There is no AI em-dash rate. There are model rates, and the human range covers most of them.
That study also turned up something better than the marker it was measuring. Asking a model for plain prose without formatting collapsed the em-dashes as a side-effect — Claude from 9.09 per thousand words to 0.19 — while asking a model directly not to use em-dashes failed to shift some of them. The habit is welded to the formatting, not to the instruction. A reader can do nothing with that, but it says something about how deep these tendencies sit: they are not preferences the system can simply be talked out of.
And since this essay was drafted with one of the models in that table, the number worth putting on the page is its own. This piece runs at 10.5 em-dashes per thousand words — 89 of them across 8,463 words of prose, counting neither tables nor matter quoted from other people, because a rate without its denominator is the thing this essay keeps objecting to. That is above the rate measured for the model that drafted it, above GPT-4.1, and more than three times the human baseline the same study reports. Which is either a confession or the argument, depending on what you think the number means. It is the argument. That range of human writers ran from 0.3 to 17, the author has written this way for thirty years, and no reader can tell those two explanations apart from the page. Neither can any detector. That is the whole point, and it is worth more here as a measurement of this essay than as a claim about anyone else’s.
Both em-dash figures come from 2026 preprints that have not been through peer review, sitting alongside work in Science Advances and PNAS that has. They are the best available on the question and they are not the same class of evidence, and nothing here should be read as though they were.
The marker that survives contact best is structural rather than lexical, and almost nobody looks for it. Machine-written prose runs noun-heavy and informationally dense, leaning on nominalisation — the trick of turning what someone did into a thing that exists — in a way that does not fit the genre it is imitating. Work in PNAS and in Scientific Reports agrees on that much, and the Scientific Reports study additionally found fewer discourse and epistemic markers in ChatGPT’s essays.
Two qualifications, since this is the marker doing the most work. The agreement between those studies is narrower than it first appears — it is firm on nominalisation and density, and thinner elsewhere, with at least one hedging-adjacent measure running the other way for one model. And nobody has tested any of it against a writer actively trying to remove it. What can be said is that it is dearer to remove than any word on a banned list, because it is a property of how the thing is built rather than of what it is called.
Then there is the layer that circulates hardest and holds up worst.
Burstiness — the variation in sentence complexity — is a real idea in statistical language processing dating back decades, wrapped around a detection claim that its own vendor abandoned. The company that popularised it stopped using it for detection in 2023 and now lists it as one indicator among several rather than as the classifier.
Invisible characters are the most persistent one, and mostly wrong. The two providers who have said anything about it say something incompatible with it. Anthropic states of Claude’s watermark, in writing, that “nothing is added to the text and there are no hidden characters”, and Google’s published scheme works on word choice rather than on characters. The reason is structural — any character-based mark dies to one find-and-replace, so no serious designer uses one.
Here is a specimen. A LinkedIn post copied and pasted into this research carried, in its timestamp line, ten visible letters interleaved with ten invisible ones — every one of them U+034F, a combining grapheme joiner. Anyone running the popular “hidden characters mean a machine wrote it” check over that paste would get a confident hit. What they had found was LinkedIn’s anti-scraping obfuscation, sitting inside a post written by a person.
AI has a smaller vocabulary is measured, and where it has been measured it runs the other way. Set against essays by students writing in a second language, ChatGPT scored higher on all three diversity measures used — on a small sample, fifty essays a side, with no statistical test reported. And the direction depends on which model you ask: the same comparison run against an earlier version of ChatGPT put the machine lower. A sign whose direction flips between model generations, and which is decided by whom you compare against, is not measuring what it claims to.
There is a fourth kind, and it is the only one an ordinary reader can see unaided. It is not a property of the writing at all. It is a trace of how the text travelled.
A post appears in a feed reading like this:
**The council has known about this for three years and done nothing. Three years.**
The pair of asterisks at each end is Markdown — the notation for bold text. It renders as bold in a chat window, in a document editor, in anything that speaks Markdown. It does not render on most social platforms, which is why it is sitting there in plain view.
Nobody typing into a post box puts asterisks around a sentence they want emphasised. Those characters arrive because the text was composed somewhere else and moved across intact. The same family includes hash marks left in front of headings, numbered lists that have lost their formatting, stray citation brackets, and the occasional instruction to the model left at the top of the post.
What it shows is less than it first appears and more useful than most of the list above: the text was written elsewhere, in a tool that speaks Markdown, and pasted without being read back. A chat assistant is the common case. A person drafting in a Markdown editor is another, and there are more of those than there used to be.
So it is not proof of a machine. It is fairly good evidence of a pipeline — and a pipeline that nobody proofread, which is its own kind of information about how much care went into the thing you are reading.
Two properties make it more useful than anything measured. It does not misfire on people writing in a second language, because it has nothing to do with how anyone writes. And it cannot be prompted away: no instruction to the model removes it, because the model is not what puts it there. It is defeated only by the person noticing before they hit post — which is why it survives at all, and why nobody has counted it.
Nobody has. There is no study of unrendered Markdown in social posts, no baseline, no rate. It goes in the list as an observation with a mechanism behind it rather than a measurement, and it is stronger than several things that do have numbers.
The asterisks have a large family. None of it has been measured either, and all of it works the same way:
| What you see | What it shows |
|---|---|
| Chat-interface furniture in a screenshot — “Claude can make mistakes”, a Regenerate button, a conversation title | An interface was on screen and was copied |
Role labels left in the text — user,
assistant, system |
Text came through a transcript or an export |
Export syntax — conversation_id,
create_time, escaped \n, mojibake like
’ where an apostrophe should be |
Text passed through a structured export, or an encoding that went wrong on the way |
HTML leakage — , &,
stray <strong> or <br> |
A rich-text copy route, or a conversion that failed |
| Answer-shaped openings — “Here is a revised version”, “Based on the information you provided”, “let me know if you’d like” | Text was lifted out of a reply to somebody |
Automation fields — {{output}}, {{date}},
a webhook fragment, a “generated by” footer |
A posting pipeline assembled it |
| Table and citation damage — flattened columns, citation numbers pointing nowhere | Text was moved between applications |
And now the distinction the whole family turns on, because it is not the one people assume.
Every item in that table is strong evidence about a transfer path and weak evidence about AI involvement. A chat panel in a screenshot shows an interface was open. It does not show that the words on it came from the model rather than from the person who pasted them in. Export syntax shows text went through software. Encoding damage shows it crossed applications. HTML leakage happens to anyone copying out of a web page, a word processor or an email.
What these marks establish is how the text travelled, which is a genuinely different question from who or what composed it. They are the strongest marks in the essay at the first question and among the weakest at the second.
Which is still worth something. A post assembled through a pipeline nobody proofread is a post whose claims deserve checking, whatever produced the sentences.
There is one source on this subject that nobody thinks to consult. In preparing this essay, several models were asked what their own output does that an outside researcher would not think to look for. The answers were consistent, unflattering, and concerned structure rather than vocabulary.
Every one of those is a discourse-layer habit, which is the layer that is expensive to remove. And every one of them also describes a competent professional writer working to a template, which is the reason none of it is proof and the reason it belongs here anyway.
None of it has been measured. It is a model’s account of itself, which is a peculiar kind of evidence, and the reader is entitled to weigh it accordingly.
Set out as a working list. The marks themselves are the least interesting column. What matters is which layer each one belongs to, what it costs to remove, and who it lands on when it is wrong.
| Sign | Layer | Evidence | What it costs to remove | What it fires on when it is wrong |
|---|---|---|---|---|
| Vendor watermark | Provenance | Strong; Google since 2024, Anthropic since Aug 2026 | Cheap: repeated paraphrasing takes detection under 10%. Translation varies by scheme | Nothing — but you cannot read it, so it cannot help you |
| C2PA credential | Provenance | Strong when present | Free: a screenshot | Nothing. Absence means nothing |
| Excess vocabulary — delves, underscores, showcasing, pivotal, realm | Language | Strong at population level; explicitly not per-text | Trivial: one instruction, or find-and-replace | Anyone who has read a lot of recent writing. It is spreading into unscripted human speech |
| Em-dashes | Language | Real at population level, useless per-text | Trivial, and by accident: a plain-prose instruction took Claude from 9.09 per thousand words to 0.19 | Every writer who likes em-dashes. Their human range covers most model rates |
| “Not X, but Y” | Language | Counted in AI output; no published human baseline found | Cheap, but has to be named specifically | Unknown, because the comparison was never run |
| Low perplexity — how predictable each next word is | Language | Measured — but measures the wrong thing | Cheap: ask for richer language | People writing in a second language, at a documented and severe rate |
| Burstiness | Language | Real statistical idea; the detection claim was dropped by its own vendor | n/a | Unknown |
| Sentence-length uniformity | Language | Direction measured, no effect size published anywhere | Trivial | Careful editing |
| Missing hedges, noun-heavy density | Discourse | Strong, replicated | Expensive: needs real rewriting | Formal registers. Technical writing. Anyone taught to write impersonally |
| Regular explanatory structure — every point developed to the same length | Discourse | Asserted | Moderate: needs restructuring, not a word swap | Anyone taught to write to a template. Professional and academic registers |
| Attention distributed evenly — argument and counter-argument at matching length | Discourse | Model self-report, unmeasured | Expensive | Careful writers. Policy and academic registers |
| No visible false starts — stable topic path, even density throughout | Discourse | Model self-report, unmeasured | Expensive | Anyone who edits before posting |
| Markdown scaffolding — bullets, bold lead-ins, tidy headings | Formatting | Asserted, no human baseline | Trivial: one instruction | Anyone writing for the web since about 2005 |
Unrendered Markdown — **bold**, stray ###,
orphaned list numbering |
Transfer | No study exists. Mechanism is plain, false-positive profile is narrow | Free, but only the person can do it: no prompt removes it | People who draft in Markdown editors. Shows a pipeline, not a machine |
| Answer-shaped opening — “Here is a revised version”, “let me know if you’d like” | Transfer | Asserted | Trivial: delete it | Tutors, consultants, support staff, anyone who writes helpfully for a living |
| Chat furniture, role labels, export syntax, HTML leakage | Transfer | Asserted; strong about the path, weak about authorship | Trivial | Anyone copying out of any web application |
| Invisible characters | Transfer | Folklore as an AI marker; no provider does this | Free | Anti-scraping systems, typography, ordinary encoding |
The second column is the one to read down, and what it shows is uncomfortable for the whole genre. Six of these marks sit in the language layer — and in the sign-lists circulating online, language is very nearly all there is, because language is what a reader notices. A list built from one layer cannot accumulate, however long it gets. The layers that would make a cluster mean something are the ones nobody collects.
Read the fourth column instead and a second pattern appears:
Almost everything cheap to detect is cheap to remove, and the marks with any staying power are the ones nobody has yet tested against someone trying to remove them.
One instruction takes a model’s em-dash rate to nearly zero. A single pass through a paraphrasing tool drops one leading statistical detector from 70.3% to 4.6%. Watermarks hold up better and still give way: one paraphrase takes detection from 99.8% to 80.7%, and repeated paraphrasing takes it from 99.3% to 9.7%.
Translation is the uneven one, and the loose version of this claim circulates widely. A round trip through another language and back takes one watermarking scheme from 94% down to 82.5%, and guts another from 91% to 26.3% — it depends entirely on which scheme, and nobody reading a post knows which. Writing in another language first and translating afterwards is the reliable attack: it drops all of them to around a fifth.
The theoretical result underneath all of it is unforgiving, though it should be stated the way its authors state it. As machine text approaches human text in distribution, the ceiling on what any detector could achieve comes down with it — and reliable detection becomes unachievable well before detection becomes worthless. The claim is not that everything ends at a coin flip. It is that the good outcome stops being available, which for anybody deciding what to believe about a stranger’s post amounts to the same thing.
Which means a sign-list is at its weakest against exactly the person you would most want to catch. It works on the person who did not think to hide it — which is generally the person doing nothing wrong.
Three reasons, and the third is the one that ought to govern how anybody uses the list above.
Signs are not independent of each other. Writes at length, writes formally, repeats themselves, will not concede, sounds angry — that looks like five findings stacking up. It is one fact seen from five angles: a person arguing hard and not much minding how it lands. Counting them as five produces a confidence nothing underneath supports, and that arithmetic is what turns a suspicion into a certainty.
Rarity beats accuracy — and be careful about what is rare. The figures at the top of this essay are shares of a corpus: two in five long-form LinkedIn posts, half of new web articles. The base rate that governs an accusation is a different number, because the accusation is about a person. What matters is how common machine-written text is among the posts you personally read and might challenge, which is a mixed feed, mostly replies and short posts, and far below those headline shares — and if that feed is one of the surfaces nobody has counted, a number nobody has. On that base, a test that catches most of what it looks for still returns mostly innocent people, because there are so many more innocent people to catch. A better test does help — but far less than intuition suggests, and only through the one property nobody advertises. What rescues a test at low prevalence is not catching more of what it is looking for; it is almost never misfiring on everything else. Nothing in this essay comes close to that, and the marks people actually use misfire constantly.
And the signs land hardest on the wrong people. Seven detectors were run over 91 TOEFL essays by people writing English as a second language. They flagged 61% of them as machine-written. On 88 essays by US eighth-graders, the same detectors flagged around 5%. The sample is small and the effect is enormous.
The mechanism is the part to carry away. When the researchers asked GPT-4 to rewrite the same non-native essays to “sound more like a native speaker”, the false-positive rate fell from 61% to under 12%. Running it backwards — simplifying the native-speaker essays — pushed their misclassification from 5% to 57%. (That last pair is printed in the preprint; the published paper gives the same result as a chart rather than as numbers.)
So the thing being measured was never machine authorship. It was how restricted the writer’s English is. Every sign built on that family is partly a proxy for the same thing, and a list that does not say so becomes a weapon aimed at people writing in a second language.
Meanwhile, people are poor at this unaided. Across 4,600 participants judging 7,600 texts, accuracy came out at 50–52% — chance. Worse, the cues participants trusted were the wrong ones: they treated first-person pronouns and mentions of personal experience as proof of humanity, which is precisely what a model produces on request.
All of which could be read as saying the marks are worthless. That would be the wrong conclusion, and the reason turns on the first of those three objections rather than against it.
Signs that are not independent cannot be added up. But it follows from the same argument that signs which genuinely are independent behave differently — and the difference is not a matter of arithmetic. It is a matter of what a single ordinary explanation can cover.
Four vocabulary items from the recognised list are not four signs. They are one sign counted four times: one register, one prompt, one writer’s habitual voice, one genre convention. Nothing has accumulated.
The same trap sits among the handling marks, and it is better hidden. A post carrying a chat-interface screenshot, role labels, export syntax and an answer-shaped opening looks like four separate discoveries. It is one action: somebody copied one reply out of one interface and pasted it without looking. Four marks, one event.
The layer column in the table is a rough guide to where a mark lives. It is not a test, and the difference matters, because the obvious next move is to make it one — count the layers, and treat four layers as four findings. That move fails, and it fails as the error the rest of this essay exists to avoid — arriving in better clothes.
Start with the counterexamples, because they are not exotic.
A corporate style guide prescribes a preferred and banned word list, an impersonal register with the point led rather than hedged, and a post template of headline, three bullets and a closing line. One document. Three layers. Entirely ordinary.
A scheduling tool used to cross-post produces both the tidy parallel formatting of its composer and the paste damage that appears wherever the destination platform’s parser differs. One tool, two layers, by mechanism rather than by coincidence.
A human translator working with standard translation software — no model involved anywhere — produces flattened vocabulary and lowered perplexity, normalised structure with the source language’s hedging conventions stripped out, and segment-level list and tag corruption. One workflow, three or four layers. And note who that lands on.
Accessibility and plain-language guidance asks for short uniform sentences, explicit heading and bullet structure, one idea per paragraph, and figurative language removed. One guideline set, three layers, firing on people writing for disabled readers and on disabled writers following authoring guidance.
None of these is an edge case. The corpus where the measurements say this is most concentrated is LinkedIn long-form — which is to say professionals posting under a house style through a scheduling tool. The layer-counting rule is least reliable exactly where the phenomenon is thickest.
And the essay’s own table gives the game away. Read down the last column for three rows in different layers — impersonal register, template-shaped structure, tidy web formatting — and they name one population: people trained to write to a house style. Three layers, one cause, printed in the table two sections ago.
Worse, my own worked example breaks the rule. Writes at length, writes formally, repeats themselves, will not concede, sounds angry. I used that as the definitive case of signs that cannot be added together. Sort those five by layer and they span at least two. Under a layer-counting rule they would qualify as a cluster — and they are, as I said then, one person arguing hard, seen from five angles.
So layers are a bad proxy for causes. The criterion was always the one stated above and it should not have been swapped for something easier to count: what a single ordinary explanation can cover.
Which leaves a rule that is more work and worth more:
Before a cluster means anything, go looking for the one pipeline that would produce all of it. A house style. A template. A scheduling tool. A translation workflow. Accessibility guidance. A single copy-paste. If one of those fits the whole cluster, you have found the explanation, and the cluster is one observation wearing several coats.
Only when nothing ordinary covers the whole set does a cluster carry more than any single mark — and even then it carries the same modest thing, which is a reason to read attentively.
One last property, and it is the one that should temper any enthusiasm for clusters. A test that requires several marks is defeated by removing the cheapest one. Read the cost column: trivial, trivial, free. Requiring four marks makes the thing easier to evade for anyone who is trying, and leaves it firing on the person who was not — who is, as the previous section says, generally the person doing nothing wrong.
Clusters are worth noticing. They are not worth building anything on.
And it cuts against the way sign-lists are normally used. The lists circulating online are almost entirely language-layer, because language is what people notice. A list of twenty vocabulary tells is one layer twenty times over, and it will produce confident wrong answers all day.
A word about the word.
“AI detection” carries a question with two answers — human or machine — and that question has been the wrong one for some time. The 71.7% figure from earlier is the plainest evidence: seven pages in ten are mixtures, running the whole distance from a light edit to a machine draft with a person’s name on the end. Text gets drafted, translated, restructured, tidied, summarised and formatted through these systems, and arrives having passed through several hands, some of which are not hands.
So the marks above do not detect anything. They indicate that an AI system may have been involved somewhere in how a post was made — which is a different claim, and one the evidence can carry. Involvement covers drafting, but it also covers translating, reformatting and packaging, and a reader has no way to tell those apart and no particular need to.
Losing the binary loses nothing worth keeping. It was never the interesting question, and the accusations that follow from it were always the accusations of somebody who had asked it.
Not a verdict. A reason to be careful.
The governing rule, stated so it can be held to:
A mark of possible AI involvement is not evidence that anyone deceived anybody, and it cannot support a claim of misconduct. It does one thing: it makes it reasonable to read a post more attentively — to weigh its claims, its sources and its apparent authority with a little more care than you otherwise would.
That is a smaller claim than the sign-lists circulating online make, and it is the only one the evidence supports. A mark on the page tells you to hold what you are reading a little more loosely — to check a claim before repeating it, to notice that forty posts agreeing might be fewer than forty judgements, to ask rather than to conclude. It does not tell you who you are talking to, and anyone who says it does has skipped every result above.
What changes is how you read, not what you conclude about the writer. Verify a factual claim rather than passing it on. Look for the primary source behind a confident summary. Notice the difference between sounding authoritative and demonstrating expertise. And stop treating polish and completeness as signals of reliability, which they were never very good evidence of even before any of this.
None of that is squeamishness. It follows from the arithmetic. A sign that fires on 61% of second-language writers cannot convict anyone of anything, and a sign that disappears when someone types one extra sentence into a prompt cannot catch anyone who cares to avoid it. Caution is what remains, and caution is genuinely useful.
Someone saying so.
There is nothing wrong with using a machine to write. It drafts, translates, tidies, and clears the way for people who find a blank page difficult or who are working in a language they learned late. The 71.7% of web pages that are some mixture of person and machine are not, in themselves, a scandal — mixture is the ordinary shape of work now, and the fact that the bucket also holds pages with barely a person in them is an argument for saying which, not for condemning the lot.
What is wrong is not saying. The problem was never that a machine was involved; it is a reader forming a view about how many people think something, when the count was manufactured. That is an arithmetical harm, and it does not require automation — a person running three accounts by hand does the same thing. Automation only makes it cheap.
And a declaration is the one instrument here that works. Consider what it survives that detection does not: it costs nothing, it needs no key, it cannot be defeated by a paraphrase, it does not misfire on anyone writing in a second language, and it does not require a single authority to sit in judgement over what is authentic. It is also the only one that reaches where the measurements never did: it works identically on an open platform and inside an encrypted thread, because it asks nobody outside the conversation to read anything.
With one condition, and without it the whole argument turns back into an accusation with better manners. The absence of a declaration means nothing. It is exactly the discipline applied to content credentials earlier: present, it tells you something; absent, it tells you nothing at all. Most people have never heard of any of this. Most posts that involved no machine will carry no note saying so, because there was nothing to say.
A norm is worth having because of what it makes available, not because of what its absence permits you to infer. The moment “they didn’t declare” becomes evidence, the marks have been handed back the job the last four sections took off them.
That last point is not a small one, and it is where the technical picture and the political one meet. Ramez Naam has argued that AI will be plural rather than singular — many models, many makers, open weights closing the gap on proprietary systems — and that this plurality is good for human freedom. On the question he is answering, he is right, and it has a consequence for everything above. A universal detector would require the very thing plurality prevents: one key, one authority, one arbiter of what is real. The reason no such detector exists is the same reason we should not want one.
Separate the question plurality answers from the one it does not, because the two get run together and the second is the one this essay is about.
Plurality answers a question about power. Many makers means no single company decides what may be said, no single key defines what is authentic, no one place to apply pressure. That is worth having, and the alternative is worse.
It does not answer a question about sameness. And the evidence for that is the strongest marker in this essay.
Consider what it means that a single list of a few hundred words identifies machine writing across vendors. If the models were plural in any deep sense, delve would mark one model rather than the category. It does not. The same excess vocabulary, the same missing hedges, the same evenly-distributed structure turn up across companies that compete hard and share almost nothing commercially. The marker works precisely because the plurality is shallow.
There is real variation at the surface. Em-dash rates run from roughly ten per thousand words in one model to zero in another, and house styles can be tuned. But underneath, these systems are built the same way, trained on overlapping copies of the same web, tuned against similar benchmarks and similar notions of a good answer. Increasingly they are trained on each other’s output, which pulls them closer rather than further apart.
An average taken over a wider base of averages is still an average. It is smoother, better sourced, harder to catch out — and narrower than the material it was drawn from, because that is what averaging does. Many models converging on a common register is not diversity of mind. It is one register with several logos on it, and its authority is greater than any single system’s would be, because it arrives from everywhere at once.
Which points at a harm this essay has not named until now, and which matters more than deception.
The problem with a great deal of machine-written text in public argument is not that it lies. Most of it does not. It is that it sounds like everything else, and it is crowding out the range of ways a thing can be said — the odd construction, the argument that runs hot for a paragraph, the point made badly by someone who has thought about it for years. Those are not stylistic decorations. They carry information about who is speaking and how much they have at stake, and a register that smooths them away removes that information while sounding more reliable for having done so.
The leakage measured earlier is the part to sit with. Delve and showcase rising in unscripted human speech, across 737,000 hours of it, means the averaging is no longer confined to the machines. People are drifting towards the mean as well, having read enough of it. A marker that leaks into human writing is a marker on its way to being useless — and the reason it becomes useless is that the thing it marked is becoming ordinary.
So plurality of vendors is necessary and it is not sufficient. It protects against one company owning the answer. It does nothing about everybody arriving at the same one.
The marks will stay unreadable, and on the surfaces where most people write there was never a mark to read. That is the price of a landscape nobody controls. Which leaves the answer that works whoever made the model, including the one running on somebody’s laptop, and needs no permission from any of them.
And it does something the detectors never could, which is the argument for it. A declaration does not tell you a post is false or that its author set out to mislead. It tells you what you are reading, so you can weigh it — and it leaves the range of human argument intact, because nobody has to write defensively to avoid being accused.
Say so. Take the answer. Argue with the argument.
It is a low standard. It is also higher than anything else on offer.
Alongside: questions and answers · glossary · sources and evidence · slides · Democratic AI°