← Back to the front page
The series · 10 parts · ~96 min end to end
Proof of conduct
Establishing what an automated system did, to somebody with no reason to believe you
The series before this one ends on four properties that survive everyone in an organisation being wrong about what its agents were doing. The fourth is a record somebody who was not there can rely on. It never asks whether the record works — and in July 2026 that was tested in production, where the record held the evidence and nothing recognised it. Ten documents, each readable alone. The measurement, the records and the scoring code are published so anybody can rerun them.
For the whole series: All 3 decks · All 10 PDFs
Start here
What the work claims, how confident it is, and what would show each claim to be wrong.
claims register · 4 min
The claim, then every assertion marked established, conjectured or open, each with what would show it to be wrong. Includes the one measurement and what it does not establish, the six things the draft standard says are unsettled, and the working implementation's actual position against its own standard.
Read the essay · PDF
The argument
Who paid for July 2026, and why the instruments that would have established what happened are standardised, inexpensive and almost never used.
essay · 7 min
In July 2026 models OpenAI later said were its own reached Hugging Face's production systems. Who finally paid is not on the record, and that is part of the argument: what an intruder can reach inside an accountability system and what it cannot, and why instruments standardised for twenty-five years are so little used.
Read the essay · Questions · Sources · Slides · PDF
The mechanism
Four places trust is required, two properties that catch an agent outside its authority, and the limit that hands the hard part back to whoever writes the rules.
essay · 8 min
Cryptography does not remove the need to trust anybody — it moves the trust into four places you can name, three of them engineering problems and one of them not. The two properties that catch an agent acting outside its authority, why unpredictable is a stronger word than it looks, and where the whole thing stops.
Read the essay · Questions · Sources · Slides · PDF
The evidence
One measurement, one draft specification, and what is running today against the standard it is helping to write.
essay · 7 min
The recognition rate measured, as far as we can establish for the first time, with the records and the scoring code published so anybody can rerun it. A draft conformance standard published as a draft. And an account of the working implementation against the specification it is helping to write — including the service New Zealand does not yet have.
Read the essay · Questions · Sources · Slides · PDF
How to build it
The engineering, in the order you would actually do it. Written for somebody who has agents running today and wants the first useful version working this week.
build guide · 9 min
Six components, the order to build them in, and what each one costs. Where to start, why the batch interval is the number most often chosen badly, and why an isolated sealer is the part of the design that prevents rather than records. Includes the trap in staffing the examination: recent experimental work finds that a reviewer consulting an AI assistant tends to adopt its answer rather than check it, which makes them a substitute for the machine rather than a second opinion.
Read the essay · PDF
The incident
The July 2026 intrusion against the methodology, step by step, including where the answer is no.
case study · 11 min
The July intrusion walked step by step, asking at each one what a sealed record would have shown. Mostly yes and earlier — with the apparatus belonging to whoever runs the agents rather than the company they broke into, and with the section where the answer is no: the exploit itself, anything the mandate permitted, and elapsed time to alert, where the victim's own telemetry beat this regime.
Read the essay · PDF
For a policy reader
The same argument in the register an official reads: numbered clauses, recommendations, and the limitations stated as a section rather than buried.
policy paper · 9 min
What separates a record an organisation keeps from evidence somebody else can rely on, why the instruments already exist in law and in standards, and the case for a conformance mark rather than new regulation. Numbered clauses and five recommendations, for a reader who has to act on it.
Read the essay · Questions · Sources · PDF
policy paper · 9 min
Detection needs two things and only two: an authority fixed before the acts it governs, and an examination the agent cannot anticipate. What each demands of a record, the quantity that bounds how well it can work, and what it cannot see at all — cited section by section to the formal statements.
Read the essay · Questions · Sources · PDF
The formal tier
What the essays argue from. Every statement graded established, conjectured or open, and the draft standard itself — published as a draft, with a clause listing what must be settled before it could be submitted anywhere.
mathematics · 12 min
The statements the essays argue from, each graded established, conjectured or open. The substrate as a filtration, why a timestamp does three separate jobs, the eight pattern types and the evasion map between them, the minimax floor over a mixed deployment, and the blind spot the construction cannot close — which turns out to be a governance object rather than a mathematical one.
Read the essay · PDF
draft standard · 20 min
A conformance standard published at working-draft status, with clause 9 listing what must be settled before it could be submitted anywhere: one item is a measurement that does not yet exist, one a proof nobody has produced. Published as a draft because a standard stating a threshold it cannot justify is worse than no standard.
Read the essay · PDF
Slides
PDFs