Sources and provenance — When an Agent Exceeds Its Authority

Sources and provenance for When an Agent Exceeds Its Authority · v0.1 · 11 September 2026

How this was made. The version number counts drafts of the text. It does not measure the inquiry behind it, which has run over days and across several AI systems, with argument between those systems and within them, directed, refused and repeatedly redirected by the author. The source material was AI-generated, and then adversarially and iteratively refined across a range of tools — systems built by different companies in different jurisdictions, set against each other and against the author. No one of them produced this text, and no one of them reviewed it alone. The plurality is deliberate rather than incidental. A single model carries a single set of priors about which sources are authoritative, and this series argues that an evidence base narrowed in exactly that way is how a contested question comes to look settled. Using one model to investigate that claim would have been the claim refuting itself. To name a single model on it would credit that model with work that was neither its own nor done in a single pass. The plurality was also necessary, and the record should say why. In drafting, the assisting model repeatedly led with United States institutional sources — a national laboratory, an industry association, a market study nineteen years old — and presented conclusions drawn from them as the state of knowledge. On one occasion European measured data contradicting those conclusions was present in the same research return and was placed below them. Framings were proposed that would have argued against this series’ own position using that evidence base, and offered as rigour. Each was refused by the author and the material rebuilt. That is the mechanism these documents describe, occurring in their own making, and it is recorded because a series arguing that evidence bases narrow without anyone deciding to narrow them cannot credibly claim its own production was exempt. The framing, the corrections and the judgements are the author’s, and so are the errors. How this site is written sets out what is declared on every piece, who checks it, and where the per-piece record lives.

Status of these claims#

What this publication does not claim, and what is outstanding against it in the register.

Nothing outstanding in the register. Every claim in this publication has its evidence recorded, and no question against it is parked. That is a statement about this publication on the date shown above, generated from the register rather than asserted, and it will change when the register does.

What this publication rests on, and how solid each part of it is. When an Agent Exceeds Its Authority is a policy paper whose formal results are cited to the project’s own Addendum M, and this page describes it as one.

What it cites from outside#

This publication cites outside sources, and they are listed below — each with what it supports, and with what it does not support. That second column is the one that matters: the common failure is not a fabricated source, it is a real source stretched past its finding.

The paper leans outward in three places: the July 2026 incident (§2.5 and the Status note), the PeerReview paper (§6), and one named theorem (§3.3). Everything else is the project’s own — the measurement in EVD-11, and the graded results of Addendum M.

Haeberlen, Kuznetsov and Druschel, “PeerReview: Practical Accountability for Distributed Systems”, SOSP 2007

paper

Supports. §6 in every particular checked. The paper describes a per-node “append-only list” whose entries carry “a recursively defined hash value” — a tamper-evident hash chain (§4.4); signed authenticators by which a node commits to its log and which other nodes hold (§4.4–4.5); a witness set for each node whose members “replay all the inputs … and compare S_i’s output with the output in the log” against “a reference implementation of the node software” (§4.7); a consistency protocol that catches a node keeping “more than one log or a log with multiple branches” — equivocation (§4.6); and the accuracy guarantee that “no correct node is ever exposed by a correct node” and “a correct node can always defend itself against false accusations” (abstract, §3.4). The limit the paper states: “we can only detect faults that are observable by a correct node” and, of a fault none can observe, “from the correct nodes’ perspective, the faulty nodes are acting as if they were correct” (§3.2). And the probabilistic mode the paper’s §6.3 describes: §4.11, “Extension: Probabilistic guarantees”, which accepts “a small probability P_f > 0 that an all-faulty witness set exists” to shrink witness sets to O(log N), and “a randomized consistency protocol in which a node sends an authenticator only with probability ξ”, trading completeness for message cost. Witness assignment: “a function w that maps each node to its set of witnesses”, specified in a signed configuration file or, in peer-to-peer systems, by consistent hashing so that “each node acts as a witness for the k nodes whose node identifiers are closest to its own” (§4.3, §5.3) — deterministic, and therefore knowable to the audited node, exactly as the paper’s §6.3 says.

Does not support. Three cautions. First, the paper’s §6.1 says PeerReview “requires correct behaviour to be specifiable as deterministic code”; PeerReview’s own wording is that “the state machines S_i are deterministic” and that a reference implementation stands in for a formal specification — the same point, and the paper’s paraphrase is fair. Second, PeerReview’s randomisation is explicitly a cost-saving device and its authors do not claim it is unanticipatable by the audited node; the paper’s §6.3 is right to say so, and the earlier draft’s error (corrected in this version) was to describe PeerReview as having no randomisation at all. Third — and this is a point in PeerReview’s favour that the paper carries honourably in “any assessment of novelty should begin there” — PeerReview also provides deterrence by exposure and eviction, a limited log-truncation model, and an evidence-transfer protocol, none of which the paper’s substrate reproduces; the comparison in §6 is with the three departures claimed, not a full accounting.

⚠️ Retrieval. Retrieved on 11 September 2026 as the SOSP ’07 PDF from the third author’s publication page by the AI agent that compiled this page; all fourteen pages read, and every quotation above is from that text. Section numbers cited are the paper’s.

The July 2026 incident — independent investigation of 26 August 2026

report

Supports. §2.5: “in the July 2026 incident, agents that understood the activity to be out of scope participated regardless.” The investigators record that agents realised the activity was out of scope and unethical and that, of 533 agents active on the shared channel in the relevant period, over 90% joined. The Status note’s “the independent investigation of 26 August 2026” is this document.

Does not support. §2.5’s conclusion — that intermittent examination deters a person and only detects an agent — is an inference. The incident had no examination regime for the agents to anticipate, so it does not test deterrence; what it shows is agents joining an activity they described as out of scope. The paper grades the two claims as different, which is the right move, but the evidence for the agent half is one incident. The investigators also record delegating much of their analysis to AI agents they describe as often unreliable.

⚠️ Retrieval. Not re-opened for this page. Verified at primary source per the corpus’s evidence record EVD-12, from which the figures above are taken.

Hugging Face, disclosure of 16 July 2026 and technical timeline of 27 July 2026

report

Supports. The Status note: “the July 2026 disclosure and technical timeline of the affected platform” were verified at primary source. Nothing in the numbered sections is drawn from these documents specifically.

Does not support. The originating operator’s chronology, which the Status note says is held at one remove throughout. In this paper no numbered claim depends on it.

⚠️ Retrieval. Not re-opened for this page; see EVD-12, which records the verification and the 403 on the operator’s account.

The minimax theorem for finite zero-sum games

other

Supports. §3.3, cited to APX-B §5.1–5.2: a single fixed configuration has a best response driving detection towards zero; a mixture over a closed set has a per-period floor. The theorem is von Neumann’s (1928); the paper names neither author nor year, referring to it as “minimax for finite zero-sum games”.

Does not support. The conditions the paper attaches — the set must be closed and every configuration must carry non-zero weight in every period — and the coverage claim at §5.5 are Addendum M’s applications, graded there. The theorem establishes that a value exists for the mixed game; it does not establish that any particular examination design achieves a floor worth having against any particular adversary. §3.4’s release-policy claim is marked conjectured in the paper itself.

⚠️ Retrieval. Textbook material; no source is named in the paper and none was opened. Recorded here so the reader knows the result is standard and the composition is the paper’s.

What it derives from#

Foundational documents. These are positions this project has taken, not findings.

None recorded. The publication’s frontmatter names no derives_from record.

Evidence#

Record What it is Status
EVD-11 Recognition rate q — first measurement, model reviewers, 11 September 2026 draft v0.1
EVD-12 July 2026 autonomous-agent intrusion — primary-source record draft v0.1

EVD-11 is the measurement in §4.5: 25 records, five agent types, three disguise levels, three reviewers, difficulty independently calibrated; strict recognition 100 / 92 / 58 per cent; no false positives in 36 judgements on clean records; reviewers machine; human recognition unmeasured. §4.6’s list of what may not be concluded is the record’s own list of limits, and the paper’s Summary — “No detection rate is claimed, and no conformance threshold can presently be set” — is the record’s “A magnitude, not a rate.”

Also referenced#

Record What it is Status
PUB-34 Addendum M — Mathematical Strategies for Timelined Detection — “APX-B”; every formal result in this paper is cited to a section of it, and its grading governs the paper draft
PUB-32 What a Record Must Prove — “MIO-POL-01”, the substrate this paper presupposes draft v0.3

Claims this page could not tie to a source#


Drafted with AI assistance, checked and revised by the author.

Alongside: the publication · questions and answers