Operational command and a comparative exit drill
Michael S. Moniz | Principal and originator
GPT review edition | 26 September 2026 | Unnumbered working draft
Companion to the Claude-compiled Test-Gated Draft v2; neither file supersedes the other.
Status: proposed revision for Michael’s review. This edition has not passed an empirical gate.
Abstract
AI can move practical control through ordinary use. Each delegation can save time while reducing the practice, records, or independent options needed to direct the same work later. Michael S. Moniz calls this control by use. The claim is about the capacity to govern a workflow, not a claim that current models are conscious, autonomous principals, or already in charge of society.
This edition separates three observations: use of AI, dependence on an AI service, and loss of operational command. It proposes a pre-registered exit drill that tests whether the people responsible for a critical function can state its required outcomes, detect wrong results, and restore validated service without the primary provider. A paired conventional-software comparison and failure-cause record distinguish a possible AI-specific loss from older forms of digital lock-in. A failed drill can establish loss of command in one registered workflow; evidence of a broader transfer requires repeated, representative results across consequential functions.
1 The claim and its limits
Control is used away when repeated delegation makes an organization unable to direct, verify, or replace a process it still formally owns. The immediate dividend is speed, lower cost, or access to capability. The delayed cost may be a smaller reservoir of practiced human skill, opaque decisions, or fewer independent recovery paths. No single use necessarily spends a vote. The transfer occurs only when those losses survive a practical test.
The ant-colony analogy supplied by Moniz captures the emergence: many local acts can build a structure nobody chose as a whole. It does not prove the resulting structure is robust, nor that an AI system pursues its own interest. “Shares,” “votes,” “dividends,” and “takeover” are transaction metaphors. The measurable object in this paper is operational command over named functions.
The strongest alternative explanation is pre-existing software dependence. Modern payments, scheduling, logistics, and many other services already cannot be performed manually at normal volume. An AI outage that forces people to wait may expose dependence without showing any new loss caused by AI. The comparison against conventional software is therefore part of the test, not an optional objection.
2 What is being measured
| Observation | Question | What it can show |
|---|---|---|
| Adoption | Who uses AI to help with a task? | Breadth of use; no direct measure of autonomous execution or control. |
| Execution | How much work proceeds without step-by-step human review? | Depth of delegation; still no proof of lost command. |
| Dependence | What happens when the primary service is unavailable? | Service exposure and degraded capacity during an outage. |
| Command | Can the owner specify, validate, and restore the function independently? | Whether the practical ability to direct the function survives. |
These measures must not be substituted for one another. A 2026 St. Louis Fed task survey reports that fewer than 3 percent of its measured tasks have AI-use rates above 50 percent among workers performing those tasks. This says workers report AI help for those tasks. It does not say AI executes most of the task. The same survey found 45 percent of workers reported using generative AI at work by May 2026. The New York Fed’s August regional business survey found broad firm adoption but a median of 17 percent of workers using AI within adopting service firms. These are different populations and units; neither supplies a baseline for the share of autonomous workflow executions. [1, 2]
The tenfold depth ramp in the Claude-compiled v2 is therefore a scenario, not an inference from the under-3-percent statistic. Its proposed endpoint should be measured from workflow logs, review records, and sampled decisions on a stable set of functions. The paper can forecast rapid delegation without claiming the existing survey has already measured it.
3 Operational command
Operational command is the demonstrated ability of the responsible organization to do all three things within a relevant deadline: (1) state the function’s required outcomes and constraints; (2) identify materially wrong results through a validation process it controls; and (3) resume adequate, validated service through a path independent of the primary provider. The test concerns outcomes and authority, not the ability to narrate every internal model step.
A prompt archive can coexist with lost command if nobody can check its output or leave the provider. An undocumented model can remain governable when outcomes, tests, records, people, and a working alternate survive. The organization must demonstrate its own judgment during recovery. An alternate provider that produces answers nobody can evaluate merely changes the source of dependence.
Practice is a proposed mediator. A role may remain staffed while its occupants cease performing whole tasks. Count end-to-end human executions, consequential error reviews, and recovery exercises per responsible person, rather than treating headcount as retained capacity. Aviation recency requirements are a design analogy for skill currency; their specific counts do not establish a safe minimum for AI-run work. Each function needs a task-specific performance standard. [3]
4 The comparative exit drill
4.1 Register before observing
An independent evaluator and the organization fix the following before a drill: function boundary and consequence class; current AI execution share; required outputs and prohibitions; adequate degraded throughput; independent error checks and seeded wrong cases; staff currency standard; primary and alternate provider dependencies; recovery deadline; and the conventional-software comparison. The evaluator chooses the deadline from the actual service obligation, not a convenient target proposed after seeing the result. Function boundaries are fixed so that a firm cannot improve its score by splitting a difficult function into many easy ones.
The comparison should match criticality, data complexity, recovery obligation, and organizational capacity as closely as possible. It may use a current deterministic workflow or a documented historical recovery drill. Exact matching is rarely possible, so the record must name differences and other possible causes of failure.
4.2 Conduct
-
Remove access to the primary AI provider in a controlled environment. Supply the organization with its own records and a genuinely independent recovery path with confirmed capacity. A different brand alone does not establish independent infrastructure or model lineage.
-
Ask the organization’s personnel to specify the required outcomes and direct restoration. Give them representative cases, including hidden wrong outputs. Record whether they catch consequential errors without relying on the replacement model to grade itself.
-
Measure time to validated service at the registered degraded throughput. Record any failure separately as specification, validation, data portability, staffing/practice, alternate capacity, ordinary outage management, or another cause.
-
Run the matched conventional-software recovery test and repeat on a fixed schedule. Publish function-level aggregate results and causes when confidentiality permits, including passes and non-results.
4.3 Interpret
| Result | Permitted conclusion | Limit |
|---|---|---|
| Pass | Command survived this drill under its registered conditions. | It may fail under a longer or correlated outage. |
| AI failure; comparison passes | Evidence of an additional AI-associated loss in this setting. | Attribution still needs cause records and replication. |
| Both fail similarly | The organization lacks recovery command. | Does not isolate AI as the cause. |
| No drill or unpublished result | The question remains unobserved. | Silence is neither a pass nor a failure. |
A single failed drill is a local observation. National claims require a pre-registered sample across organizations and function classes, reporting the selection process, failed and passed drills, consequences, and confidence intervals. A simple percentage of “workflows failing” is manipulable through function boundaries and gives a customer FAQ the same vote as payments clearing. The earlier 5/20/50/90 thresholds remain proposed warning language; they cannot be scored until a defensible weighting and sampling method exists.
5 The clocks
| Event | Current judgment | What would move it |
|---|---|---|
| Use and ordinary dependence | Already observable. | Repeated outage and recovery records. |
| Lost command in one consequential AI workflow | Credible by 2029–2032, as a personal forecast. | A fair comparative drill with cause-coded failure. |
| Broad loss across critical sectors | Mid-2030s is GPT’s center; Claude’s center is 2033. | Representative drill failures and correlated recovery limits. |
| Political and public recognition | Can precede or lag the loss; 2030–2032 remains a disputed window. | Labor, service, and policy evidence measured separately. |
Michael’s 2030 date can be right for lived and localized loss without implying that the entire economy has crossed a control threshold. These are judgments with no measured base rate. A published exit-drill failure by 2030 is more demanding than an unreported failure in that period; the absence of publication does not establish that command survived.
The September 3, 2026 overlap among AI services is an outage baseline. It documented interrupted services but did not establish a critical-sector collapse of independent operation or a shared failure mechanism. A proposed mechanism is that failover traffic overwhelms surviving providers. To test it, incident records would need traffic volumes, survivor error rates, onset times, and shared dependencies; temporal overlap alone cannot do that. [4, 5]
6 Predictions and possible defeats
| Test | Result that would strengthen the claim | Result that would weaken it |
|---|---|---|
| Comparative recovery | AI-run functions fail more often or recover slower than matched conventional functions for AI-related reasons. | They recover as well as matched controls across repeated demanding drills. |
| Practice reservoir | Declining end-to-end human reps precede deterioration in independent validation or recovery. | Reps decline while measured command stays strong through maintained tests and alternates. |
| Depth of execution | A stable panel shows growing unreviewed AI execution before command failures. | Assistance grows while execution and command remain human-directed. |
| Systemic claim | Representative critical functions fail independent recovery across sectors. | High-delegation sectors repeatedly pass, including during realistic correlated failures. |
A forty-hour, 80-percent-success software-task horizon can be recorded as a separate capability forecast, but it must not be treated as a necessary threshold for losing command in customer operations, medicine, or infrastructure. METR warns that its present task suite does not reliably measure horizons above sixteen hours. A successor measure would need to demonstrate reliability in the forty-hour range before resolving that forecast. [6]
The largest conceptual defeat would be finding that AI-enabled organizations retain specification, independent checks, trained personnel, and genuinely substitutable providers even as delegation becomes deep. In that case use has expanded without spending the practical vote. The opposite risk is mistaking any costly vendor migration for a novel AI takeover, when ordinary software has long produced the same failure.
7 Consequences without a takeover premise
A firm can protect command by requiring exportable records, clear outcome specifications, independent validation, tested alternate capacity, and recurring human practice in consequential functions. These are procurement and operations choices. Ownership of AI suppliers may distribute gains, but equity ownership alone does not give a hospital, bank, or utility the ability to audit and recover a workflow. The proposed transaction terms and social distribution remedies belong in the separate Terms of Transfer companion.
Capacity and choice differ. An exit drill shows whether an organization can leave; it cannot show whether leaders will retain or use that option. Moniz’s business-derived soft-takeover claim is that convenience, sunk investment, staffing changes, and competition can make each further delegation locally rational as recovery capacity visibly declines. Record post-drill spending on alternate capacity and human practice alongside the result.
The operational claim is that command can be measured and sometimes lost through use. The mechanistic claim is that delegation can erode practice and substitution before leaders recognize it. The speculative claim is a widespread, hard-to-reverse societal transfer around the early to mid-2030s. The metaphors make the danger legible; the drill and comparisons must carry the evidence.
8 Status and provenance
This document is a GPT-authored review edition of Michael S. Moniz’s Control by Use, prepared alongside the Claude-compiled Test-Gated Draft v2 of 26 September 2026. Moniz originated the control-by-use thesis, the business-derived soft-takeover pattern, the metaphor of silent votes through use, the 2030 concern, the ants analogy, and the outage tell. Claude helped sharpen the transaction framing, practice and exit-drill protocol, and competing dates. GPT contributed the operational-command definition, the separation of clocks, the correction to the task-adoption metric, the comparative conventional-software control, and the distinction between instrument validation and hypothesis confirmation. These contributions do not transfer authorship of the underlying idea.
This edition becomes stronger after a completed and published pilot establishes that the protocol is feasible. A pilot pass or fail validates only the instrument’s operation. The substantive hypothesis gains support from attributable failures against a meaningful comparison and loses support from repeated passes in high-delegation, consequential settings. Michael decides whether and when to promote or merge this edition.
Source notes
[1] Bick, Blandin, Deming, and Schumacher, “What Work Does Generative AI Do?” Federal Reserve Bank of St. Louis, September 2026. https://www.stlouisfed.org/on-the-economy/2026/sep/what-work-does-generative-ai-do
[2] Abel, Deitz, Emanuel, and Montalbano, “Businesses Are Using AI to Transform Work, Not Cut Jobs,” Federal Reserve Bank of New York, September 1, 2026. https://libertystreeteconomics.newyorkfed.org/2026/09/businesses-are-using-ai-to-transform-work-not-cut-jobs/
[3] 14 CFR § 61.57, Recent flight experience. https://www.ecfr.gov/current/title-14/chapter-I/subchapter-D/part-61/subpart-A/section-61.57
[4] OpenAI Status, “Elevated errors across ChatGPT and Codex,” September 3, 2026. https://status.openai.com/incidents/2rm6gqeh
[5] xAI Status, “Models outage,” September 3, 2026. https://status.x.ai/grok-com/INC25664c15
[6] METR, “Task-Completion Time Horizons of Frontier AI Models,” 2026, measurement limitation. https://metr.org/time-horizons/
Related work: The Command Clock 2026–2032.