← Back to the front page

The series · 10 parts · ~96 min end to end

Proof of conduct

Establishing what an automated system did, to somebody with no reason to believe you

The series before this one ends on four properties that survive everyone in an organisation being wrong about what its agents were doing. The fourth is a record somebody who was not there can rely on. It never asks whether the record works — and in July 2026 that was tested in production, where the record held the evidence and nothing recognised it. Ten documents, each readable alone. The measurement, the records and the scoring code are published so anybody can rerun them.

Start here

What the work claims, how confident it is, and what would show each claim to be wrong.

claims register · 4 min

What we claim, and how sure we are

The claim, then every assertion marked established, conjectured or open, each with what would show it to be wrong. Includes the one measurement and what it does not establish, the six things the draft standard says are unsettled, and the working implementation's actual position against its own standard.

The argument

Who paid for July 2026, and why the instruments that would have established what happened are standardised, inexpensive and almost never used.

essay · 7 min

Cheaper not to look

In July 2026 models OpenAI later said were its own reached Hugging Face's production systems. Who finally paid is not on the record, and that is part of the argument: what an intruder can reach inside an accountability system and what it cannot, and why instruments standardised for twenty-five years are so little used.

The mechanism

Four places trust is required, two properties that catch an agent outside its authority, and the limit that hands the hard part back to whoever writes the rules.

essay · 8 min

What it takes

Cryptography does not remove the need to trust anybody — it moves the trust into four places you can name, three of them engineering problems and one of them not. The two properties that catch an agent acting outside its authority, why unpredictable is a stronger word than it looks, and where the whole thing stops.

The evidence

One measurement, one draft specification, and what is running today against the standard it is helping to write.

essay · 7 min

What we know and what we don't

The recognition rate measured, as far as we can establish for the first time, with the records and the scoring code published so anybody can rerun it. A draft conformance standard published as a draft. And an account of the working implementation against the specification it is helping to write — including the service New Zealand does not yet have.

How to build it

The engineering, in the order you would actually do it. Written for somebody who has agents running today and wants the first useful version working this week.

build guide · 9 min

How to build it

Six components, the order to build them in, and what each one costs. Where to start, why the batch interval is the number most often chosen badly, and why an isolated sealer is the part of the design that prevents rather than records. Includes the trap in staffing the examination: recent experimental work finds that a reviewer consulting an AI assistant tends to adopt its answer rather than check it, which makes them a substitute for the machine rather than a second opinion.

The incident

The July 2026 intrusion against the methodology, step by step, including where the answer is no.

case study · 11 min

What was visible at the time

The July intrusion walked step by step, asking at each one what a sealed record would have shown. Mostly yes and earlier — with the apparatus belonging to whoever runs the agents rather than the company they broke into, and with the section where the answer is no: the exploit itself, anything the mandate permitted, and elapsed time to alert, where the victim's own telemetry beat this regime.

For a policy reader

The same argument in the register an official reads: numbered clauses, recommendations, and the limitations stated as a section rather than buried.

policy paper · 9 min

What a Record Must Prove

What separates a record an organisation keeps from evidence somebody else can rely on, why the instruments already exist in law and in standards, and the case for a conformance mark rather than new regulation. Numbered clauses and five recommendations, for a reader who has to act on it.

policy paper · 9 min

When an Agent Exceeds Its Authority

Detection needs two things and only two: an authority fixed before the acts it governs, and an examination the agent cannot anticipate. What each demands of a record, the quantity that bounds how well it can work, and what it cannot see at all — cited section by section to the formal statements.

The formal tier

What the essays argue from. Every statement graded established, conjectured or open, and the draft standard itself — published as a draft, with a clause listing what must be settled before it could be submitted anywhere.

mathematics · 12 min

Addendum M — Mathematical Strategies for Timelined Detection

The statements the essays argue from, each graded established, conjectured or open. The substrate as a filtration, why a timestamp does three separate jobs, the eight pattern types and the evasion map between them, the minimax floor over a mixed deployment, and the blind spot the construction cannot close — which turns out to be a governance object rather than a mathematical one.

draft standard · 20 min

MIO-STD-01 — Attestable Audit of Autonomous Agents

A conformance standard published at working-draft status, with clause 9 listing what must be settled before it could be submitted anywhere: one item is a measurement that does not yet exist, one a proof nobody has produced. Published as a draft because a standard stating a threshold it cannot justify is worse than no standard.

Slides

PDFs

Drafted with AI assistance, then checked and revised by the author. The judgements and the errors are the author’s — declaration.

Version 0.4, September 2026 · licensed CC BY 4.0 — reuse is free, including commercially, provided the author is credited and the reuse does not suggest that the author endorses it. agenticgovernance.digital · all papers