The terms the essay uses, in plain words, with what each is worth.
Democratic AI°. Used here in a defined sense, marked with a degree sign. Not AI made in a democracy, not AI sold by firms headquartered in democracies, not AI that is cheap or widely available. A measure of whether the people a system affects can govern it, contest it, refuse it, and retain authority over the knowledge it is built from.
⚠️ Not a claim about any country. The measure is about architecture, and it returns the same answer whichever jurisdiction a system comes from. The essay quotes an American submission because it is the clearest written example of the phrase being used to mean ours; the open-weight models most organisations deploy are increasingly Chinese, and they fail the fourth condition for identical reasons.
The degree mark (°). Signals that the four conditions measure a position on a scale rather than award a badge. Treated as a binary category the term would have almost no members. As a degree it returns two of four, and here is which. It is not a trademark and claims no ownership of the phrase.
The four questions. Govern · contest · refuse · retain authority. The whole instrument.
Plurality. How many suppliers there are. A claim about market structure. Measurable by counting firms.
Epistemic independence. Whether those suppliers give you genuinely separate evidence, separate ways of being wrong, separate ideas of what a good answer looks like. Not measurable by counting firms — which is the essay’s whole point.
False corroboration. Asking several systems the same question, getting the same answer, and treating that as confirmation. It feels like triangulation. It is closer to asking one source repeatedly, because the systems share data, tuning norms and incentives.
Correlated error. Two systems being wrong in the same way rather than in different ways. The measurable version of the point above, and the strongest evidence in the essay.
Baseline (or chance rate). How often something would happen by luck alone. Without it a percentage is uninterpretable, and it depends on the setup: two guessers picking among three wrong options agree 33% of the time by accident, so a measured 60% is 1.8 times chance rather than near-total agreement. Where there are more options the baseline falls — 12.7% on the other dataset in the same study — and the same measured gap becomes a bigger effect.
Effect size. How much bigger the finding is than the baseline. The reason 42.3% against a 12.7% baseline (3.3×) is a stronger result than 60% against 33% (1.8×), despite being the smaller headline number.
Peer-reviewed. Assessed by independent specialists before publication. Not a guarantee of truth; a floor.
Preprint. Posted publicly before that assessment. Often good, sometimes wrong, and marked as such throughout because a reader is entitled to weigh it differently.
Null result. A study that looked for an effect and did not find it. Worth more than it usually gets — except where its authors say their instruments were unsuited to the question, which is the case with the one null result here.
Selection cascade. The narrowing from everything people know, through what is written, digitised, crawlable, licensed, filtered and preference-tuned, to what a model will say. Every stage excludes something. Note the limit: the same shape describes print publishing, which narrowed harder. Selection is not the same as convergence.
Preference tuning. Training a model after the fact on human judgements of which answers are better. It produces the helpful, balanced, complete register — and that register is a choice, not a neutral default.
Benchmark culture. The shared set of tests by which systems are ranked. It does not prove models are identical. It does mean everyone is rewarded for the same capability profile — a common answer to “what counts as a better system?”
Distillation. Training a smaller model on a larger one’s output. Ordinary practice, and one reason distinct-looking systems converge.
Model collapse. What happens when models are trained repeatedly on their own generated output: rare material at the edges of the distribution drops away. Demonstrated in Nature — and contested, because the experiments replaced real data with synthetic rather than letting real data accumulate alongside it. The essay treats it as a reason for vigilance, not a prediction.
Provenance. Where a claim came from and who stands behind it. The essay’s argument is that compression removes this before it removes content — a model can preserve what was said while deleting who was entitled to say it, and a similarity score will report that nothing was lost.
Register. The characteristic sound of a kind of writing. The concern is not that the machine register is bad, but that it can become the standard by which other speech is judged — an unacknowledged admission test for credibility.
Situated knowledge. Knowledge that cannot be cleanly separated from its context: who holds it, what obligations attach, what authority is needed to pass it on. It travels badly as decontextualised text, which is what these pipelines are built to handle.
Māori data governance. The principle that Māori exercise authority over Māori data. Te Kāhui Raraunga, acting for the Data Iwi Leaders Group, publishes the current operative instruments — the Māori Data Governance Model (Tuia te korowai o Hine-Raraunga) and the Māori AI Governance Framework, whose core statement is that AI systems must not be implemented in Aotearoa without fully realising Māori authority over Māori data. Te Mana Raraunga, the Māori Data Sovereignty Network, remains active as an advocacy body; ⚠️ its 2016–2018 principles predate AI and, per Dr Karaitiana Taiuru’s September 2025 critique, do not address model training or algorithmic discrimination. Related: the CARE Principles (Global Indigenous Data Alliance), and WAI 262 — a Treaty of Waitangi claim and Tribunal report, not a data governance framework.
Open weights. A model whose parameters are published, so anyone can run, inspect or adapt it. The strongest part of the plural-AI case — and the same property that makes any universal marking or disclosure regime unenforceable. Freedom and ungovernability described by people with different worries.
Compute concentration. That the visible plurality of assistants rests on a much narrower base: over 90% of notable models from industry, over 60% of global compute from one supplier, almost all leading chips from one fabricator.
Chokepoint. A layer with few enough participants that control there is control over everything above it. Diversity at the layer people check can sit on concentration at the layer that decides who can build anything.
Designed plurality. Independence built deliberately across data, models, evaluation, infrastructure, institutions, interfaces and public culture — rather than assumed from the number of vendors.
Correlated-failure testing. The practical version: give candidate systems the same hard cases, record which items each gets wrong, and check whether the failures overlap more than chance predicts. The one recommendation here that an organisation can act on without waiting for anybody.
Redundancy in name only. Buying three systems, believing you have three fallbacks, and reproducing one blind spot through all three.
Falsifiability. Naming in advance what would show you were wrong. The essay lists six such conditions, one of which is close to satisfied — which is the point of writing them down.
Alongside: the essay · questions and answers · sources and evidence · slides · The Marks It Leaves