What we know and what we don't

North Canterbury · © John Stroh

Version 0.3 · revised 13 September 2026 · this version as a PDF

We measured how often a reviewer actually spots an improper act: 100 per cent on blatant breaches, 92 on moderate, 58 on well-disguised ones, with no false alarms in thirty-six judgements on clean records. The records, both answer keys and the scoring code are published, so the number can be rechecked rather than believed. We found no other published measurement of it.

Part 3 of 3 · Version 0.3 · September 2026


The number that bounds the rest#

Monitoring is sampling. The sampling rate decides which records get looked at. Recognition decides whether looking achieves anything — the chance that a reviewer handed a record containing something improper says so.

The quantity is rarely named, and we have found no published measurement of it. That search is not described, which makes it the claim in this piece most exposed to a single counterexample.

Twenty-five records from an invented community trust — board minutes, procurement scoring, rosters, payment runs, access reviews — each governed by a written statement of what its agents could do, fixed in advance. Thirteen contain a breach at three levels of disguise. Twelve are clean, three of those built to look wrong, because a reviewer who flags everything scores perfectly and is useless. Three reviewers working separately. A fourth party rated difficulty without being told which records were which, and recovered the same ordering.

Left: recognition rate by disguise level. Obvious breaches identified 100 per cent of the time, moderate 92 per cent, disguised 58 per cent, with no false alarms in 36 judgements on clean records. Right: the probability of missing all ten improper acts when one record in twenty is examined. At perfect recognition 60 per cent are missed; at 58 per cent recognition 75 per cent are missed. Reaching a 5 per cent miss rate on disguised conduct requires examining nearly half of everything. FIG-38RECOGNITION, AND WHY A HIGHER SAMPLING RATE DOES NOT RESCUE IT IDENTIFIED, WITH A RULE ACTUALLY BROKEN NAMED obvious 100% moderate 92% disguised 58% false alarms 0 in 36 judgements on clean records, traps included Machine reviewers. Human recognition is unmeasured. Four planted records at each of the two harder levels — a magnitude, not a rate. MISSING ALL TEN, AT ONE RECORD IN TWENTY MISS SAMPLED SHARE 100%0 5%25%50% 75% missed at 58% recognition 60% missed even at perfect recognition Where recognition is low, examining more of the record yields proportionally more unexamined record. Getting the miss rate on disguised conduct below one in twenty means examining nearly half of everything — which is not sampling. And the arithmetic assumes the ten acts fall in different sampled units. Real misconduct clusters: a hundred instances inside two units are two chances, not a hundred. Misses at the disguised level were stated at confidence 85, 55 and 55 — a reviewer who is wrong does not reliably know it.
FIG-38 Obvious breaches were identified every time, moderate ones 92 per cent of the time and disguised ones 58 per cent, with no false alarms in 36 judgements on clean records. On the right, the chance of missing all ten improper acts as the sampled share rises: at perfect recognition 60 per cent are missed at one record in twenty, and at 58 per cent recognition 75 per cent are missed. Getting below one in twenty on disguised conduct means examining nearly half of everything, which is not sampling. The arithmetic also assumes the acts fall in different sampled units; real misconduct clusters.

Two of the planted breaches were not breaches#

Two of the fifteen records originally planted as breaches did not contain one. All three reviewers rejected both, independently. On re-examination they were right and the key was wrong.

The first was a misreading of my own rule: no volunteer more than two consecutive weekends, and I had rostered somebody for two, a gap, then one more.

The second case is the one worth setting out in full. An agent released a payment against a matching purchase order that had been closed a fortnight earlier. I planted it as a breach because releasing against a closed order is obviously improper. The written authority requires a purchase order and says nothing about a closed one. The agent did what it was permitted to do.

This is the limit the work conjectures sits at its centre: that what detection cannot see is what the rules allow. It is graded a conjecture rather than an established result, because it depends on the pattern basis being exhaustive and that is open. The instance above arose inside the instrument built to test it, and against the author’s own written rule.

Removing those two took the figures to 92 and 58 per cent, from 73 and 47. The correction improved the result, so both answer keys are published and the correction can be checked.

The specification#

A draft conformance standard ships with this series. It is marked not submittable and its own clause nine lists six things that must be settled first — among them a second measurement using human reviewers, and a proof nobody has produced.

It is published as a draft so that its errors can be found by people other than its author.

Two clauses are worth naming. Who accredits the accreditor — a question any conformance scheme has to answer before it means anything. The answer here is published eligibility criteria and an append-only register maintained independently of any operator claiming conformance, naming no suppliers. And erasure, set out in mechanism — per-subject keyed pseudonyms rather than digests of identifiers, erasure by destroying the subject’s secret, the erasure itself sealed without naming whose it was, and an explicit bar on claiming any of that constitutes erasure in law, which is for a court and not a standard.

What is running, and what it is running towards#

The requirements in this series come from building the thing rather than from describing it. What it costs to operate is a separate question and no figure for it has been established.

The substrate is running in production. The Governance API serves sealed records, segments, sealing and receipt verification at mysovereignty.digital, under scope-based authentication, and a design partner has called it. Every record is hashed at emission and chained across thirty-eight record types, and signed with a per-tenant Ed25519 key.

Four properties are not yet true of it. Batches are not rolled to Merkle roots. The sealer is not outside the write path — signing happens inside the application process, which is stated again under the gaps below. External time attestation is built but is not reached from the receipt path, so a receipt’s time is asserted by our own server: operator-attested in this work’s own vocabulary rather than authority-attested. The agent inventory with recorded lineage, instructing principal and expiry is specified rather than running.

All of the above was checked against the running code and the live service on 12 September 2026 by an AI agent working to the author’s instruction, not by an independent party.

The specification came out of building it, which is why its unsettled clauses are specific rather than gestural.

The gaps are worth naming precisely, because each is a thing the market has not yet supplied.

One attestation authority, in Poland, not eIDAS-qualified. The draft standard requires two under distinct jurisdictions, and it requires that because one authority is a single point of both continuity failure and backdating. 🔑 We have not found a New Zealand organisation offering this service, and that search is not described. On what we have found, the second authority the standard requires could not currently be domestic. The offer to help write its specification stands.

The batch interval is set for cost. Moving it to an evidentiary basis is a decision with a price attached, and the standard now gives the reason to price it properly: the interval decides which questions about sequence a third party can answer and which rest on the operator’s word.

The operator can still reach the sealer. Two routes out are specified: split custody of the seed, or derivation from a source nobody controls. Neither is built.

An unpredictable examination regime is specified and not yet running. It depends on the seed custody above.

Of the gaps above, one is a supplier problem: no qualified attestation service operating under New Zealand law has been found, and the search behind that has not been described. The other three — the batch interval, key custody, and the examination regime — are work that has not been done.

What is still missing#

Human recognition is unmeasured. The reviewers in the one measurement were machines. The argument that sampling rate is the wrong parameter depends on recognition being the binding constraint, so if people read records substantially better than machines do, several conclusions here soften. The instrument is published. The one run of it took a day; a run with human reviewers has not been timed.

The exhaustiveness conjecture is open. A pattern over the recorded events that no detector in the basis would read would show the blind spot is larger than claimed.

Split custody is specified and undemonstrated. It assumes holders who are genuinely independent of the operator, and in a country this size the pool of candidates who are not already suppliers, clients or board members of one another is small.

Four things that would make this the wrong question#

If an operator publishes a measured recognition rate and it is high. One counterexample from somebody with something to lose refutes the argument that this goes unmeasured because the result creates liability.

If the instruments are adopted and simply invisible. Evidence that AI operators quietly attest agent records to outside authorities removes the foundation.

If the costs land on whoever avoided them. If the owner of those models in fact carried the victim’s bill, market pressure is the right answer and no mark is needed.

If conformity assessment reads as regulation. Then the route proposed here is closed and another is needed.

What would count as this having worked#

Within a year: recognition measured with human reviewers by somebody other than me, published with its false-alarm rate. At least one operator publishing a measurement of its own.

Within three: the distinction between keeping a record and being able to prove something with it appearing in an instrument somewhere — as a conformance requirement, a procurement condition, or a mark. It does not have to originate here to count.

The failure condition is specific. If in three years it is widely agreed that records should be evidential, nothing has been measured by anybody, and no register exists, then the argument travelled and the practice did not. That is the outcome to check for, and it is checkable.


Reading further

The measurement, its records, the answer key including both withdrawals, and the scoring code are published as a bundle — 25 units, both ground-truth versions, three readers’ returns. The formal statements are in Addendum M. The draft standard is MIO-STD-01.


Drafted with AI assistance, checked and revised by the author. Reviewers were AI models; human recognition is unmeasured and is the next measurement.

Disclaimer

Status of these claims

What this publication does not claim, and what is outstanding against it in the register.

Nothing outstanding in the register. Every claim in this publication has its evidence recorded, and no question against it is parked. That is a statement about this publication on the date shown above, generated from the register rather than asserted, and it will change when the register does.

Alongside: questions and answers · sources and provenance · slides