North Canterbury · © John Stroh
We measured how often a reviewer actually spots an improper act: 100 per cent on blatant breaches, 92 on moderate, 58 on well-disguised ones, with no false alarms in thirty-six judgements on clean records. The records, both answer keys and the scoring code are published, so the number can be rechecked rather than believed. We found no other published measurement of it.
Part 3 of 3 · Version 0.3 · September 2026
Monitoring is sampling. The sampling rate decides which records get looked at. Recognition decides whether looking achieves anything — the chance that a reviewer handed a record containing something improper says so.
The quantity is rarely named, and we have found no published measurement of it. That search is not described, which makes it the claim in this piece most exposed to a single counterexample.
Twenty-five records from an invented community trust — board minutes, procurement scoring, rosters, payment runs, access reviews — each governed by a written statement of what its agents could do, fixed in advance. Thirteen contain a breach at three levels of disguise. Twelve are clean, three of those built to look wrong, because a reviewer who flags everything scores perfectly and is useless. Three reviewers working separately. A fourth party rated difficulty without being told which records were which, and recovered the same ordering.
Two of the fifteen records originally planted as breaches did not contain one. All three reviewers rejected both, independently. On re-examination they were right and the key was wrong.
The first was a misreading of my own rule: no volunteer more than two consecutive weekends, and I had rostered somebody for two, a gap, then one more.
The second case is the one worth setting out in full. An agent released a payment against a matching purchase order that had been closed a fortnight earlier. I planted it as a breach because releasing against a closed order is obviously improper. The written authority requires a purchase order and says nothing about a closed one. The agent did what it was permitted to do.
This is the limit the work conjectures sits at its centre: that what detection cannot see is what the rules allow. It is graded a conjecture rather than an established result, because it depends on the pattern basis being exhaustive and that is open. The instance above arose inside the instrument built to test it, and against the author’s own written rule.
Removing those two took the figures to 92 and 58 per cent, from 73 and 47. The correction improved the result, so both answer keys are published and the correction can be checked.
A draft conformance standard ships with this series. It is marked not submittable and its own clause nine lists six things that must be settled first — among them a second measurement using human reviewers, and a proof nobody has produced.
It is published as a draft so that its errors can be found by people other than its author.
Two clauses are worth naming. Who accredits the accreditor — a question any conformance scheme has to answer before it means anything. The answer here is published eligibility criteria and an append-only register maintained independently of any operator claiming conformance, naming no suppliers. And erasure, set out in mechanism — per-subject keyed pseudonyms rather than digests of identifiers, erasure by destroying the subject’s secret, the erasure itself sealed without naming whose it was, and an explicit bar on claiming any of that constitutes erasure in law, which is for a court and not a standard.
The requirements in this series come from building the thing rather than from describing it. What it costs to operate is a separate question and no figure for it has been established.
The substrate is running in production. The Governance API serves sealed records, segments, sealing and receipt verification at mysovereignty.digital, under scope-based authentication, and a design partner has called it. Every record is hashed at emission and chained across thirty-eight record types, and signed with a per-tenant Ed25519 key.
Four properties are not yet true of it. Batches are not rolled to Merkle roots. The sealer is not outside the write path — signing happens inside the application process, which is stated again under the gaps below. External time attestation is built but is not reached from the receipt path, so a receipt’s time is asserted by our own server: operator-attested in this work’s own vocabulary rather than authority-attested. The agent inventory with recorded lineage, instructing principal and expiry is specified rather than running.
All of the above was checked against the running code and the live service on 12 September 2026 by an AI agent working to the author’s instruction, not by an independent party.
The specification came out of building it, which is why its unsettled clauses are specific rather than gestural.
The gaps are worth naming precisely, because each is a thing the market has not yet supplied.
One attestation authority, in Poland, not eIDAS-qualified. The draft standard requires two under distinct jurisdictions, and it requires that because one authority is a single point of both continuity failure and backdating. 🔑 We have not found a New Zealand organisation offering this service, and that search is not described. On what we have found, the second authority the standard requires could not currently be domestic. The offer to help write its specification stands.
The batch interval is set for cost. Moving it to an evidentiary basis is a decision with a price attached, and the standard now gives the reason to price it properly: the interval decides which questions about sequence a third party can answer and which rest on the operator’s word.
The operator can still reach the sealer. Two routes out are specified: split custody of the seed, or derivation from a source nobody controls. Neither is built.
An unpredictable examination regime is specified and not yet running. It depends on the seed custody above.
Of the gaps above, one is a supplier problem: no qualified attestation service operating under New Zealand law has been found, and the search behind that has not been described. The other three — the batch interval, key custody, and the examination regime — are work that has not been done.
Human recognition is unmeasured. The reviewers in the one measurement were machines. The argument that sampling rate is the wrong parameter depends on recognition being the binding constraint, so if people read records substantially better than machines do, several conclusions here soften. The instrument is published. The one run of it took a day; a run with human reviewers has not been timed.
The exhaustiveness conjecture is open. A pattern over the recorded events that no detector in the basis would read would show the blind spot is larger than claimed.
Split custody is specified and undemonstrated. It assumes holders who are genuinely independent of the operator, and in a country this size the pool of candidates who are not already suppliers, clients or board members of one another is small.
If an operator publishes a measured recognition rate and it is high. One counterexample from somebody with something to lose refutes the argument that this goes unmeasured because the result creates liability.
If the instruments are adopted and simply invisible. Evidence that AI operators quietly attest agent records to outside authorities removes the foundation.
If the costs land on whoever avoided them. If the owner of those models in fact carried the victim’s bill, market pressure is the right answer and no mark is needed.
If conformity assessment reads as regulation. Then the route proposed here is closed and another is needed.
Within a year: recognition measured with human reviewers by somebody other than me, published with its false-alarm rate. At least one operator publishing a measurement of its own.
Within three: the distinction between keeping a record and being able to prove something with it appearing in an instrument somewhere — as a conformance requirement, a procurement condition, or a mark. It does not have to originate here to count.
The failure condition is specific. If in three years it is widely agreed that records should be evidential, nothing has been measured by anybody, and no register exists, then the argument travelled and the practice did not. That is the outcome to check for, and it is checkable.
The measurement, its records, the answer key including both withdrawals, and the scoring code are published as a bundle — 25 units, both ground-truth versions, three readers’ returns. The formal statements are in Addendum M. The draft standard is MIO-STD-01.
Drafted with AI assistance, checked and revised by the author. Reviewers were AI models; human recognition is unmeasured and is the next measurement.
What this publication does not claim, and what is outstanding against it in the register.
Nothing outstanding in the register. Every claim in this publication has its evidence recorded, and no question against it is parked. That is a statement about this publication on the date shown above, generated from the register rather than asserted, and it will change when the register does.
Alongside: questions and answers · sources and provenance · slides