Questions and answers — What we know and what we don't

Questions and answers for What we know and what we don't · v0.2 · 13 September 2026

How this was made. The version number counts drafts of the text. It does not measure the inquiry behind it, which has run over days and across several AI systems, with argument between those systems and within them, directed, refused and repeatedly redirected by the author. The source material was AI-generated, and then adversarially and iteratively refined across a range of tools — systems built by different companies in different jurisdictions, set against each other and against the author. No one of them produced this text, and no one of them reviewed it alone. The plurality is deliberate rather than incidental. A single model carries a single set of priors about which sources are authoritative, and this series argues that an evidence base narrowed in exactly that way is how a contested question comes to look settled. Using one model to investigate that claim would have been the claim refuting itself. To name a single model on it would credit that model with work that was neither its own nor done in a single pass. The plurality was also necessary, and the record should say why. In drafting, the assisting model repeatedly led with United States institutional sources — a national laboratory, an industry association, a market study nineteen years old — and presented conclusions drawn from them as the state of knowledge. On one occasion European measured data contradicting those conclusions was present in the same research return and was placed below them. Framings were proposed that would have argued against this series’ own position using that evidence base, and offered as rigour. Each was refused by the author and the material rebuilt. That is the mechanism these documents describe, occurring in their own making, and it is recorded because a series arguing that evidence bases narrow without anyone deciding to narrow them cannot credibly claim its own production was exempt. The framing, the corrections and the judgements are the author’s, and so are the errors. How this site is written sets out what is declared on every piece, who checks it, and where the per-piece record lives.


The reviewers were AI models and you say no threshold can be set from the result. Why publish it?#

Because we have found no other published measurement — a search we have not described — and because the draft standard requires the quantity be measured for a deployment rather than assumed (MIO-STD-01 §8.3).

The number is a magnitude, and EVD-11 says so in those words. What it settles is that recognition is neither near one nor near zero: near-perfect on obvious breaches, roughly three in five on disguised ones, which is exactly the regime in which raising the sampling rate cannot compensate. At five per cent sampling, ten disguised instances are missed about three times in four (§9, item 1). It also reports the figure that makes a recognition rate mean anything — zero false alarms in thirty-six judgements on clean records — which §8.3 requires, because a reader that flags everything scores perfectly and is useless. The instrument is published so that the next measurement, with human reviewers, can be run by somebody else.


You changed your own answer key after scoring and your result went up. Why should anyone accept that?#

Both answer keys are published, so the change can be inspected rather than taken on trust. Two records planted as breaches were withdrawn during scoring; the figures moved from 73 and 47 per cent to 92 and 58.

The reason is that all three reviewers, working separately, rejected both records, and on examination they were right. One was a misreading of the author’s own rostering rule. The other is the central limit of the work occurring inside the instrument built to test it: an agent released a payment against a closed purchase order, the written authority required a purchase order and said nothing about a closed one, and so the act was permitted. The uncorrected run had counted the author’s two errors as reviewer misses. Both withdrawals are in the published answer key, and the draft standard now requires any future measurement to publish its ground-truth corrections the same way (§9, item 1).


Twenty-five records and four planted breaches per level. Is that a measurement at all?#

It is one measurement, and EVD-11 lists what it cannot support: model reviewers only; too few units for a confidence interval; every record seen in one context, so independence between them is imperfect; synthetic material written by the person who planted the breaches; one mandate in one setting.

What it does support is stated with equal care. A fourth party rated every record for difficulty without knowing which were breaches and recovered the authored ordering — mean 1.40, 1.75 and 3.00 on a three-point scale, every disguised record rated hard — so the levels are not an artefact of the author’s framing. The piece names what would make it more than one measurement: human reviewers, somebody other than the author running it, and a false-alarm rate published alongside. Within a year, it says.


Your implementation uses one attestation authority, in Poland, and your own standard requires two. Is it failing its own standard?#

It is a working implementation ahead of its market, and the piece describes it as a proof of concept in those terms.

The standard requires two authorities under distinct jurisdictions (§5.2, §5.3) because a single authority is a single point of both continuity failure and backdating (§5.2.2). The implementation has one, in an EU member state, not eIDAS-qualified. The reason it does not have a second one in New Zealand is that we have not found a New Zealand organisation offering the service, and that search is not described. The specification says what such a service would have to do, and the offer to help write that specification stands. The standard itself lists attestation-authority selection among the six things to settle before submission (§9, item 4).


You say the batch interval is an evidentiary parameter, and also that yours was set for cost. Which is it?#

Both, and the second is a departure from the first that the claims register records as such.

Addendum M §2.3 establishes (E3) that two events in the same batch are unordered under the external authority, so any claim about sequence finer than the batch interval rests on the operator’s own clock. The interval therefore decides which questions a third party can answer independently and which rest on the operator’s word. The implementation’s interval was chosen for token cost. The construct records that as a decision with evidentiary consequences rather than a tuning knob, and the draft standard requires the interval to be recorded and treated as evidentiary (§5.6, §5.8). What the standard adds is the reason to price a shorter interval properly, against the timescale of the patterns that must be attested independently.


The operator can still reach the sealer and holds the seed. Can the operator not tell an agent when it is watched?#

Yes. The construct’s actor map lists the operator as holding the sealer, seed, inventory and observer keys, and rates the exposure low as intent and high as capability — which is what a threat model records.

The route out is known and specified: split custody of the seed with the threshold above the number of holders the operator controls, or derivation from an external source nobody controls (MIO-STD-01 §7.2.1). Which of the two to adopt is Open 11 in the construct. The unpredictable examination regime depends on that decision and is specified rather than running.


Why publish a specification marked “not submittable”?#

Because a specification held back until its author has resolved everything is a specification whose errors are all its author’s.

Clause 9 lists six things that must be settled first and says it must not be removed to make the document look ready: the recognition rate measured once and for model readers only; the exhaustiveness conjecture open; the delta over PeerReview to be stated precisely, since PeerReview in 2007 already composed most of the substrate and stated the residual as a known limit; attestation authority selection; the intellectual-property regime, which is a decision for an editor; and alignment with ISO/IEC 24970 on AI system logging, unexamined. Two of those are a second measurement using human reviewers, and a proof nobody has produced.


An append-only record and a right to erasure cannot both hold. How does the standard handle that?#

By specifying a mechanism and declining to claim it is erasure in law.

MIO-STD-01 §4.6 bars sealing personal data in clear or encrypted form; requires that a natural person be referenced by a keyed pseudonym under a per-subject secret rather than a digest of an identifier; provides erasure by destroying that secret; requires the erasure itself to be sealed without naming whose it was; requires the residual risk to be disclosed to the subject; and bars any claim that this constitutes erasure as a matter of law, which is for a court. The policy paper in this series sets it against the European Data Protection Board’s Guidelines 02/2025 and says the distinction — sealing the shape of events and never their content — must be defended against that guidance rather than assumed to be accepted.


What does it look like if this has worked, and what if it has not?#

Within a year: recognition measured with human reviewers by somebody other than the author, published with its false-alarm rate, and at least one operator publishing a measurement of its own. Within three: the distinction between keeping a record and being able to prove something with it appearing in an instrument somewhere — a conformance requirement, a procurement condition, or a mark. It does not have to originate here to count.

The failure condition is specific. If in three years it is widely agreed that records should be evidential, nothing has been measured by anybody, and no register exists, then the argument travelled and the practice did not. That is checkable, and whoever looks back should check it.


Drafted with AI assistance, checked and revised by the author.

Disclaimer

Status of these claims

What this publication does not claim, and what is outstanding against it in the register.

Nothing outstanding in the register. Every claim in this publication has its evidence recorded, and no question against it is parked. That is a statement about this publication on the date shown above, generated from the register rather than asserted, and it will change when the register does.

Alongside: the publication · sources and provenance