Sources and provenance — How much it may do unsupervised

Sources and provenance for How much it may do unsupervised · v0.2 · 8 September 2026

How this was made. The version number counts drafts of the text. It does not measure the inquiry behind it, which has run over days and across several AI systems, with argument between those systems and within them, directed, refused and repeatedly redirected by the author. The source material was AI-generated, and then adversarially and iteratively refined across a range of tools — systems built by different companies in different jurisdictions, set against each other and against the author. No one of them produced this text, and no one of them reviewed it alone. The plurality is deliberate rather than incidental. A single model carries a single set of priors about which sources are authoritative, and this series argues that an evidence base narrowed in exactly that way is how a contested question comes to look settled. Using one model to investigate that claim would have been the claim refuting itself. To name a single model on it would credit that model with work that was neither its own nor done in a single pass. The plurality was also necessary, and the record should say why. In drafting, the assisting model repeatedly led with United States institutional sources — a national laboratory, an industry association, a market study nineteen years old — and presented conclusions drawn from them as the state of knowledge. On one occasion European measured data contradicting those conclusions was present in the same research return and was placed below them. Framings were proposed that would have argued against this series’ own position using that evidence base, and offered as rigour. Each was refused by the author and the material rebuilt. That is the mechanism these documents describe, occurring in their own making, and it is recorded because a series arguing that evidence bases narrow without anyone deciding to narrow them cannot credibly claim its own production was exempt. The framing, the corrections and the judgements are the author’s, and so are the errors. How this site is written sets out what is declared on every piece, who checks it, and where the per-piece record lives.

Status of these claims#

What this publication does not claim, and what is outstanding against it in the register.

Nothing outstanding in the register. Every claim in this publication has its evidence recorded, and no question against it is parked. That is a statement about this publication on the date shown above, generated from the register rather than asserted, and it will change when the register does.

What this publication rests on, and how solid each part of it is. How much it may do unsupervised is an essay, and this page describes it as one.

What it cites from outside#

This publication cites outside sources, and they are listed below — each with what it supports, and with what it does not support. That second column is the one that matters: the common failure is not a fabricated source, it is a real source stretched past its finding.

Feng, McDonald and Zhang, Knight First Amendment Institute essay series

paper

Supports. Five levels by role retained (operator, collaborator, consultant, approver, observer); quoted: “While an agent may be instructed to seek approval prior to taking consequential actions, reliably determining which actions are consequential may be challenging,” and “User disengagement can lower the care with which they approve actions and can result in unintended approvals” — used both as a prior autonomy scale and as evidence for the rubber-stamping failure mode.

Does not support. The document flags it is “an essay rather than a peer-reviewed paper, which is worth stating because it is frequently cited as though it were the latter” — published with a single anonymous reviewer.

Morris and colleagues, Google DeepMind (preprint)

paper

Supports. Cited as one of three existing autonomy scales — six levels from “No AI” through “AI as a Tool,” “Consultant,” “Collaborator,” “Expert” to “AI as an Agent — fully autonomous AI.”

Does not support. Described only as “a preprint, widely cited” — not peer-reviewed; used only to note existence and structure, not for any further claim.

Mitchell, Ghosh, Luccioni and Pistilli, Hugging Face

paper

Supports. Five levels ordered by program-flow control; quoted: “we argue fully autonomous AI agents… should not be developed,” because “risks to people increase with the autonomy of a system.” Cited as the one of three scales “willing to say a level should not exist,” backing this piece’s own position that Delegated Autonomy “should generally not be used” outside a sandbox.

Does not support. A research group’s normative position, not independent proof; the piece adopts the conclusion as its own stated position rather than as established fact.

SAE J3016 (driving-automation levels)

standard

Supports. Cited as “the canonical precedent” establishing that automation level is a property of task and circumstances, not a fixed attribute of the machine — the feature this piece argues the AI scales failed to carry over.

Does not support. The document did not read the standard directly and cannot confirm exact wording beyond secondary reproduction.

⚠️ Retrieval. ⚠️ “J3016 is behind a paywall; the level definitions here are as reproduced in secondary references citing it, not read from the standard.”

Bradshaw, Hoffman, Johnson and Woods, “The Seven Deadly Myths of ‘Autonomous Systems’”, IEEE Intelligent Systems, May/June 2013

paper

Supports. Quoted as “the strongest published argument against everything in this piece”: Myth one, “‘Autonomy’ is unidimensional”; Myth two, that levels-of-autonomy is useful scientific grounding for roadmaps; verdict quoted: “levels of autonomy encourage reductive thinking,” and autonomy “isn’t a discrete property of a work system… it’s an idealized characterization.”

Does not support. The document concedes the objection applies to its own tiers (“a single ordered scale, which is the thing Bradshaw and colleagues say should be discarded”) and offers only a contested reply, left explicitly open: “Whether that reply is sufficient is a fair question.”

US Defense Science Board (quoted by Bradshaw et al.)

report

Supports. Quoted via Bradshaw and colleagues for the claim that levels-of-autonomy taxonomies are misread as implying “that autonomy is simply a delegation of a complete task to a computer, that a vehicle operates at a single level of autonomy and that these levels are discrete and represent scaffolds of increasing difficulty.”

Does not support. Reached only through Bradshaw et al.’s quotation; not consulted directly by this piece.

NIST’s ALFUS framework

Supports. Cited as splitting autonomy across three axes (human independence, mission complexity, environmental complexity) “precisely to avoid collapsing it into one” — offered as more faithful to autonomy than a single ordered scale, supporting the Bradshaw objection.

Does not support. Not claimed to be adopted for AI agent governance; cited only as a contrasting multi-axis design written for unmanned systems.

EU AI Act, Article 7(2)(c),(d),(g),(h),(i)

regulation

Supports. Quoted criteria for classifying high-risk systems — data sensitivity, extent of autonomy/human override, dependence of potentially harmed persons, power imbalance/vulnerability, and reversibility — used as five of the six factors argued to move an action’s permitted autonomy level down.

Does not support. “Article 7 uses these to classify systems for regulatory purposes, whereas an institution needs them to set the level for classes of action. The criteria transfer; the unit of application does not. An organisation adopting them is not implementing the AI Act.”

Commission Delegated Regulation (EU) 2017/589 (RTS 6)

regulation

Supports. Quoted requirements that an algorithmic-trading firm be able to “cancel immediately, as an emergency measure, any or all of its unexecuted orders,” identify “which trading algorithm and which trader… is responsible for each order,” and (Article 15) implement pre-trade price collars and volume/message limits — used as precedent for bounded execution over per-action approval.

Does not support. The document itself flags in its closing section that the precedent might not transfer: orders are “homogeneous and typed” while agent actions “are neither.”

SEC Rule 15c3-5

regulation

Supports. Quoted: requires controls “reasonably designed to prevent the entry of orders that exceed appropriate pre-set credit or capital thresholds, or that appear to be erroneous,” under “the direct and exclusive control of the broker or dealer with market access” — used with RTS 6 as precedent for structural pre-set limits.

Does not support. In force for securities market access since 2010; not claimed to have been applied to AI agents.

SEC order regarding Knight Capital, 1 August 2012 incident

report

Supports. Quoted finding the firm “did not have adequate safeguards in place to limit the risks posed by its access to the markets,” with figures (“more than four million orders,” “397 million shares,” “more than $460 million” lost in 45 minutes) — the documented consequence of operating without a structural boundary; “the Commission’s first enforcement action under the market access rule.”

Does not support. A single securities-trading incident, used as illustration, not as a general AI-agent failure-rate statistic.

Article 29 Working Party, guidance on automated decision-making, adopted 2017, last revised February 2018

framework

Supports. Quoted at length: “The controller cannot avoid the Article 22 provisions by fabricating human involvement… the controller must ensure that any oversight of the decision is meaningful, rather than just a token gesture. It should be carried out by someone who has the authority and competence to change the decision” — used to document that draft-and-approve is “the level with the best-documented failure.”

Does not support. Interpretive guidance on Article 22 GDPR; does not itself address AI agents generally or non-EU jurisdictions.

Court of Justice of the European Union, SCHUFA (C-634/21), 7 December 2023

court decision

Supports. Cited holding that a credit score is itself an automated individual decision where recipients “attribute to it a determining role in the granting of credit” — used to show the human downstream “does not launder the decision if the human defers to it.”

Does not support. A ruling on credit scoring under Article 22 GDPR specifically; the analogy to AI-agent approval workflows generally is this document’s own extension, not the Court’s holding.

“one widely cited body of survey research on software delivery” (unnamed report, published by a cloud vendor; later called “the DORA finding”)

report

Supports. Cited for the finding that “formal external approval bodies were associated with worse delivery performance” and no evidence a more formal external review process reduced failures — used to support that an approval step can cost throughput “without buying safety.”

Does not support. The document restricts the claim itself: “it is published by a cloud vendor, it surveys software deployment rather than AI agents, and it is self-reported… a single cross-domain survey statistic cannot carry a conclusion about AI approval regimes.” What it supports is narrower: that approval is “worth something only where the approver has the authority, the competence and the time to refuse.”

⚠️ Retrieval. Self-reported, vendor-published, cross-domain (software delivery, not AI agents); figures deliberately not reproduced by the document.

Privacy Act 2020 (New Zealand), consolidated form

statute

Supports. Searched and reported: “The word ‘automated’ appears zero times. The word ‘algorithm’ appears once, inside the information-matching provisions” — used to establish that New Zealand has “no equivalent to Article 22, no right to human intervention in an automated decision, and no threshold at which oversight becomes mandatory”.

Does not support. A word-search finding about the current consolidated statutory text only; does not examine case law or other instruments that might separately bear on automated decisions.

Privacy Commissioner Michael Webster, statement of December 2025 (fifth anniversary of the Privacy Act)

evidence

Supports. Quoted: “We also need stronger protections for the significant privacy risks that arise from automated decision-making, which can cause problems such as inaccurate predictions, discrimination, unexplainable decisions, and a lack of accountability” — used as regulator confirmation that the protections are absent — the essay’s point being that this is “not an inference from silence. The regulator says so.”

Does not support. A public call for reform, not itself a change in law.

What it derives from#

Foundational documents. These are positions this project has taken, not findings.

Record What it is Status
CON-05 Default deny draft v0.1

Evidence#

None. This publication references no evidence record. That is the correct description of what it is rather than a gap: it is an essay, reasoning from the foundational documents above rather than reporting a measurement. Where it states a number, that number is marked in the text as what it is.

Also referenced#

Record What it is Status
CON-06 Most restrictive wins draft v0.1
CON-10 Bounded reversibility, honestly stated draft v0.1
FIG-28 What “human in the loop” conceals drawn for this publication · figures/FIG-28.svg

Generated from the corpus, not written by hand: this page cannot claim a source the corpus does not hold, and it changes when the records do.

Alongside: the publication · questions and answers