Questions and answers — Cheaper not to look

Questions and answers for Cheaper not to look · v0.2 · 13 September 2026

How this was made. The version number counts drafts of the text. It does not measure the inquiry behind it, which has run over days and across several AI systems, with argument between those systems and within them, directed, refused and repeatedly redirected by the author. The source material was AI-generated, and then adversarially and iteratively refined across a range of tools — systems built by different companies in different jurisdictions, set against each other and against the author. No one of them produced this text, and no one of them reviewed it alone. The plurality is deliberate rather than incidental. A single model carries a single set of priors about which sources are authoritative, and this series argues that an evidence base narrowed in exactly that way is how a contested question comes to look settled. Using one model to investigate that claim would have been the claim refuting itself. To name a single model on it would credit that model with work that was neither its own nor done in a single pass. The plurality was also necessary, and the record should say why. In drafting, the assisting model repeatedly led with United States institutional sources — a national laboratory, an industry association, a market study nineteen years old — and presented conclusions drawn from them as the state of knowledge. On one occasion European measured data contradicting those conclusions was present in the same research return and was placed below them. Framings were proposed that would have argued against this series’ own position using that evidence base, and offered as rigour. Each was refused by the author and the material rebuilt. That is the mechanism these documents describe, occurring in their own making, and it is recorded because a series arguing that evidence bases narrow without anyone deciding to narrow them cannot credibly claim its own production was exempt. The framing, the corrections and the judgements are the author’s, and so are the errors. How this site is written sets out what is declared on every piece, who checks it, and where the per-piece record lives.


Hugging Face’s own systems caught the intrusion. Is this not a detection failure rather than a record failure?#

It was a detection failure first, and the piece quotes the sentence that says so: the pipeline correlated the signals and then “failed to correctly raise the alert’s criticality and trigger the on-call team.” Nothing proposed here would have raised that alarm. The companion piece on the July timeline says plainly where a sealed record would have shown nothing.

The record question comes afterwards, and it is about two different sets of records. Hugging Face reconstructed the intrusion from its own security telemetry, which it held and the intruder did not. Separately, the independent investigation published in August found that roughly seven per cent of the agents’ own transcripts had been spoofed — those records belonged to the party whose agents were acting, and that is where the editing problem sits. Establishing what an agent did from the agent’s own account is the thing this series is about. It is a separate failure from the alarm that did not sound.


You say an attestation does not make a record true. So what does it buy?#

Less than the marketing vocabulary suggests, and the guarantee is narrow. An intruder inside early enough can seal days of invention, and every page will carry a perfect timestamp.

What it removes is the retrofit. Once the shape of the argument is known, nobody can go back and adjust the record to suit. Formally, the sealed substrate is a filtration: what was sealed by time t is fixed at t for all later times, and corrections are new events that reference the old (Addendum M §2.1, graded established as E1). European law attaches a presumption to exactly this property — Regulation 910/2014, Article 41(2), the accuracy of the date and the integrity of the data bound to it. That is a real thing to hold when the other party is the one whose agents wrote the record.


Append-only records date from 1991 and time attestation from 2001. If this were worth doing, would it not have been done?#

The instruments are cheap and available, and the piece reports finding no product in this class that uses them. The obstacle is not technical.

Each property that makes a record credible to an outsider removes a degree of freedom from the insider: a record that cannot be amended, a clock the organisation does not control, rules fixed in advance that somebody else can measure it against. Assembled, they make a record that can be used against you, by people you cannot choose. The two fields usually reached for — nuclear safeguards and post-crisis financial supervision — arrived after a crisis; no instrument or agency is cited for that, and Certificate Transparency is a counter-case, adopted because browsers came to require it. Nobody in the piece is accused of hiding anything. The claim is that the incentive points one way, and that price pressure does not reach it because a buyer has no way to see the difference.


Most vendors writing “immutable” and “tamper-proof” mean it. Is the accusation one of bad faith?#

No. Three different things are sold in the same words. A record is a statement you must trust the holder about. An integrity check proves the text has not changed and says nothing about when or by whom. An attestation brings in a party who was not the author. Each is a legitimate product.

The problem is that a purchaser comparing two of them has no way to tell which they are buying, so nothing in the comparison rewards the stronger one. That is why the piece asks for a mark struck by somebody other than the maker rather than for better behaviour from makers. Gold has carried such a mark since the medieval period for the same reason: the seller’s word about fineness was never the point.


The government has ruled out new AI regulation. Is this not asking for exactly that?#

The national strategy says the approach is light-touch and requires no additional regulatory overlay. Conformity assessment is not regulatory overlay; it is how a claim becomes checkable, and it is the ordinary route for regulated goods.

The standards are in place — NZS ISO/IEC 42001 and 23894 have been nationally adopted. What is missing is any statement of how anyone would be checked against them: the strategy does not contain the words assurance, audit, certification or conformity. The gap is between adopting a standard and being able to say who conforms to it, and closing it needs no statute. There is also a date: New Zealand chairs the 2027 Digital Nations meeting.


You measured your own recognition rate and got three in five on disguised breaches. Does that not argue against you?#

The measurement is 100 per cent on blatant breaches, 92 on moderate and 58 on well-disguised ones, with no false alarms in thirty-six judgements on clean records. Fifty-eight is the hardest of three levels, not the result.

The argument does not depend on any of those numbers being high. It depends on their being measured, and on our not having found another published one — a search we have not described.

The measurement is published with its records, its answer key and its scoring code, and with the figure that gives the recognition column its meaning: zero false alarms in thirty-six judgements on clean records. Two units originally planted as breaches were withdrawn during scoring because all three reviewers correctly said they were not breaches, and the correction moved the figures up (EVD-11).

The run took a day. That is what it cost to measure the quantity. It says nothing about what it would cost an organisation to operate the apparatus, and no figure for that has been established.


If neither company behaved badly and both disclosed more than they had to, why should the bill have gone anywhere else?#

Nobody is blamed in the piece, and both companies are credited for what they published. Nor does the piece say where the bill landed: who finally paid is not on the record, and neither document establishes it.

What the record does show is the structural point. The party that ran the investigation could not inspect the software that carried out the intrusion, because it was not theirs. Whether they could have prevented it is not established either way — their own pipeline did detect it. The question the piece raises is not who paid but whether the records could answer it, and they cannot.

The piece also names the observation that would overturn this. If OpenAI in fact carried Hugging Face’s bill, market pressure is doing the work and no mark is needed. OpenAI’s own account was not retrievable when this was verified and is cited at one remove (EVD-12), so the question is open in fact and the piece says so.


What would make you drop the argument?#

Four things, each stated in the piece. An operator publishing a measured recognition rate that is high — one counterexample from somebody with something to lose. Evidence that operators already attest agent records to outside authorities and simply do not say so. Evidence that the costs of the July intrusion landed on whoever avoided them. And officials reading conformity assessment as regulation, which would close the route proposed and require another.


Drafted with AI assistance, checked and revised by the author.

Disclaimer

Status of these claims

What this publication does not claim, and what is outstanding against it in the register.

Nothing outstanding in the register. Every claim in this publication has its evidence recorded, and no question against it is parked. That is a statement about this publication on the date shown above, generated from the register rather than asserted, and it will change when the register does.

Alongside: the publication · sources and provenance