Questions and Answers — Democratic AI°

Including the objections that go against the essay, and the one that nearly sank it.

The argument

Is this an argument against competition in AI?

No — the opposite. Competition constrains vendor power, reduces lock-in, lowers prices and widens access, and open weights make local deployment and inspection possible. Attempts to turn regulation into a barrier to entry deserve resisting.

The objection is narrower: plurality answers a question about market structure and is used as reassurance about knowledge and public language. Those are different claims and the second does not follow.

So what would count as real plurality?

Independent evidence paths, independent evaluation frameworks, independent error pathways, and materially different institutional grounding. Not price, interface, latency or benchmark score.

The short version:

Counting models tells you about the market. It tells you nothing about whether they know different things.

Isn’t “they all sound the same” just an impression?

It would be if that were the argument. It isn’t. The load-bearing evidence is about errors, not style: across two datasets covering 71 and 349 models, when two of them are both wrong they pick the same wrong answer far more often than chance. That is measurable, and it was measured.

The numbers

You say models agree 60% of the time when both are wrong. Isn’t that just chance?

Partly, and the essay says so rather than hoping you won’t ask.

The 60% comes from the HELM dataset: 71 models, four options per question, so three wrong ones and a chance agreement rate of 33%. The finding is 1.8 times chance.

The same study’s other dataset — the HuggingFace leaderboard, 349 models — has more options per question, so chance there is 12.7%. Measured agreement was 42.3%: 3.3 times chance, the stronger result despite being the smaller number. Quoting the 60% is quoting the weaker of the two findings.

Almost everything quoting that study gives the 60% without the 33%. In a piece arguing that compression preserves the claim and strips the conditions that make it meaningful, doing that would have been an odd choice.

Doesn’t that weaken your argument?

It narrows it, which is different. Two findings from the same paper matter more than the headline anyway:

The second is the uncomfortable one. If shared error were simply a symptom of immaturity you would expect it to fall as systems improve. It does not.

How far does a multiple-choice benchmark really carry?

Not as far as some readings want. It is the narrowest object in the field — one right answer, no register, no provenance, no interpretation. It supports the rule that agreement between models is not corroboration. On its own it establishes nothing about ways of knowing, and the essay says so.

The counter-evidence

Has anyone found that public language is actually narrowing?

Not yet, and this is the most important answer on the page.

The one naturalistic study available compared about 30,000 English news articles from 2018 with 30,000 from 2024. It found no decrease in lexical diversity — one measure went up. It also found a significant rise in AI-associated vocabulary over the same period.

So the cause is present and increasing, and the predicted effect is absent.

Doesn’t that falsify the whole thesis?

It very nearly does, and it was checked carefully for that reason. What stops it is the authors’ own limitations section:

“we suspect the lexical diversity methods we applied are inappropriate for revealing a loss of lexical diversity on the scale of a very large text corpus. Therefore, our empirical contribution to the hypothesis… remains highly limited.”

That is a measurement failure reported as one, not a clean result in either direction. It is also a single genre, a student workshop paper, with no significance tests and a word list borrowed from biomedical abstracts.

The position the evidence supports:

The mechanism is demonstrated in controlled settings. The population-scale effect is plausible, uneven, and not yet shown.

Then isn’t the essay unfalsifiable — always calling for more measurement?

That is the sharpest objection and it deserves a straight answer rather than a deflection.

The essay’s falsification list names six conditions, and one of them is close to satisfied already. The qualifier attached to it — that corpora must be measured with instruments their own authors consider fit for the purpose — is doing real work rather than protecting the thesis, because the authors in question are the ones who disowned their instruments.

If a study using metrics its authors stood behind found no homogenisation across comparable genres, that would count. Nobody has run it.

The cascade

Your selection cascade — doesn’t it prove too much?

Yes, and the essay concedes it in the section where the diagram appears.

Draw the same shape for print publishing: literacy, then a publisher, then an acquiring editor, then house style, then a distributor. It narrows harder at every stage, in fewer languages, admitting orders of magnitude fewer items. Nobody thinks twentieth-century print produced a monoculture of mind.

So selection alone proves nothing. The diagram states a mechanism; it cannot carry the conclusion. The argument rests on the measurements, which is why they are ranked by strength rather than listed as a wall.

Doesn’t AI actually widen access — more languages, more people writing?

In some respects, plainly yes, and that is a real point against the concern. Generation is cheap in languages that were previously expensive to publish in, and demand for training data has driven digitisation of archives that were not digital.

The essay’s claim is not that fewer things exist. It is that what returns through the systems has been pulled toward a common centre — and that agreement between those systems is not evidence about the world.

The uncomfortable parts

The essay is AI-drafted and argues AI writing homogenises prose. Isn’t that self-defeating?

Logically, no — an argument is not false because of who made it. But the sharper version of the attack lands, so it is worth stating.

The essay asks for source trails, separation of quotation from synthesis, and disclosure where AI materially shaped a claim. A one-line declaration does not supply what its own theory of provenance says compression removes. What it does supply is the sources page: every study named, its status marked, its limits stated, and three places where checking changed the argument.

Check the working rather than the byline. That is the whole method.

But your fact-checking was done by AI agents. Isn’t that the false corroboration you warn about?

It would be, if the agents had been checking each other. That is the sharpest objection to this piece and it deserves the essay’s own rule applied to the essay.

The rule is: the comparison that counts is not model against model, it is a claim against material you can inspect. So that is what was done. The agents did not vote on whether a claim sounded right. They fetched the papers, opened the PDFs and read the numbers in the tables — the same documents named on the sources page, which you can open yourself.

And where two agents disagreed, the disagreement was settled by which one produced a document. Twice, one reported it could not find a study that another had already located with a working identifier and a verbatim quotation. A failure to find is weaker evidence than a quotation with an identifier, and treating those two as cancelling out would have cut true and sourced claims.

Separate instances of a model are not separate sources. Primary documents are. The method works only because the checking terminated in documents rather than in agreement — and if it had not, the objection would land.

Why does the Indigenous data governance section stop where it does?

Because that failure is extraction and consent, and the essay is about convergence. They are different problems. A hundred genuinely different models, each trained without consent on restricted knowledge, commits the error a hundred times over — designed plurality does not touch it.

The essay says that plainly rather than borrowing the moral weight of an argument it does not actually support. It also names the frameworks that state the point more precisely than the essay does, because a claim made in someone else’s territory should follow them rather than paraphrase them.

Who decides what is restricted?

Not a procurement officer. For Māori data specifically, Te Kāhui Raraunga’s Māori AI Governance Framework is unambiguous: AI systems must not be implemented in Aotearoa without fully realising Māori authority over Māori data. The essay does not attempt to improve on that. The general question — who decides what is restricted, for whose knowledge — is not one this essay answers, and it should not borrow an answer given about Māori data and apply it to everybody’s. The general question — who decides what is restricted, for whose knowledge — is not one this essay answers, and it should not borrow an answer given about Māori data and apply it to everybody’s.

Using the measure

Does anything actually score four out of four?

Not that we are aware of, and saying so is the point. Treated as a badge the term would be empty — a standard nothing meets is a slogan. Treated as a degree it does useful work, because it tells you which condition a system fails and therefore what would have to change.

Then score something. Score your own.

Fair. The system this site’s authors build — Village AI, a community-governed deployment — comes out at roughly three of four.

Question Village AI
Govern — can the community govern it? Yes. That is the design
Contest — can a decision be challenged, and change? Yes
Refuse — is declining available and survivable? Yes — self-hosted, no lock-in
Retain authority over the knowledge No, at the base-model layer

The fourth fails for a reason worth stating plainly: it runs on a foundation model somebody else trained by scraping the open web — a Chinese one, as it happens, chosen for open weights and local deployment rather than for its flag. That is the point. The condition fails on the architecture, not on the jurisdiction. The people whose writing is in that corpus never agreed to it, cannot see what was taken, and cannot ask for it back. Everything built above that floor is governed properly. The floor is not.

Nobody currently passes that condition, short of training a foundation model from consented data — a nine-figure undertaking. Which is exactly why the measure returns a position rather than a verdict.

Doesn’t retrieval fix that? If the answers come from a corpus the community governs, isn’t the base model just grammar?

Partly, and it is the strongest available answer — but it moves the score rather than closing the gap, and this essay’s own argument is what limits it.

What retrieval genuinely fixes: the substance of an answer comes from a corpus that can be added to, restricted, withdrawn and cited. Federation strengthens it further — if each community governs its own corpus rather than pooling into one index, authority stays distributed by construction rather than by policy.

What it does not fix:

This essay argues that compression strips provenance before it strips content. Retrieval protects the content. The conditions of transmission and the register are the part it does not reach.

So is that a claim or a guess?

A well-founded expectation, and it is stated as one rather than asserted, because the credibility of everything else here depends on that distinction.

There is a test that would settle it. Take questions a governed corpus answers well. Run them with retrieval against one base model, then against a materially different one. If the substance holds and only the phrasing moves, retrieval is doing the work. If framing, emphasis and conclusions move too, the base model is deciding more than assumed — and you will know which half of the fourth condition has actually been secured.

Nobody has published that. It is the same shape as the shared blind spot test, and we would rather someone ran it than agreed with us about it.

What to do

What can I actually do about any of this?

If you buy AI systems for an organisation, one thing, and it is cheap:

  1. Give every candidate system the same held-out set of hard cases from your real domain
  2. Record which items each one gets wrong, not just how many
  3. Check whether the failures overlap more than chance predicts
  4. Make decorrelation a tender condition, not branding
  5. Re-run it at renewal, because correlation rises as systems improve

Three vendors is not three judgements. Without that test, multi-vendor procurement is redundancy in name only.

And if I am just a reader?

Treat agreement between systems as a prompt to look at a source, not as a source. The comparison that counts is a model’s claim against something you can inspect — a document, a dataset, a person accountable for having said it.

Should I stop using these tools, or write more roughly to prove I am human?

No. That converts an institutional problem into an aesthetic obligation imposed on individuals, which is unfair and does not work. The essay says so explicitly, and it is the reason the practical section is addressed to procurement rather than to people’s prose.

Alongside: the essay · glossary · sources and evidence · slides · The Marks It Leaves