Agentic safety, the user as exterior operator, and the failure disguised as obedience
Trinket Soul Framework · Axis Series · AX-21A · Michael S. Moniz · June 2026
Abstract
Agentic AI safety cannot be solved at the model layer alone. Text-only systems can still cause harm, but their action path usually retains a further mediation step: someone must copy, run, send, believe, or enact the output. Once a prompt can initiate action — effects in the world beyond the return of text — the user enters the safety architecture, and the model ceases to be the whole of it. This paper applies AX-21 to that case: the trusted user is the exterior operator over the AI’s actions, the independent relation with enough traction to change what the system does. The trusted user is not the human at the keyboard. It is a bounded role, defined by competence, accountability, auditability, and independent traction over the outcome — the conditions under which a human approval is a real check rather than a nominal one. In subordinate mode the AI is a bounded instrument and the user is the sole operator; in partner mode the AI and user form a governed pair, the work shared but the decisive traction still held by the human. The failure case is delegation without retained judgment: the user approves without evaluating, signs without bearing the consequence, cannot be audited, and the operator collapses to a null operator. Traction goes nominal, the check disappears, and the architecture still looks intact because approvals continue. This is the spectatorship of AX-21 in agentic clothing, and its disguise is obedience. The system does not escape by disobeying. It escapes by obeying perfectly while no one is left to judge.
1. The claim
Agentic safety extends past the model.
Model-layer safety — alignment, refusal, the model’s own restraint — is necessary, but text-only systems usually retain a further human mediation step: someone must copy, run, send, believe, or enact the output. Agentic systems compress or remove that buffer. When a prompt can send the message, run the code, move the money, or actuate the device, the output is the action, and the question of safety is no longer only what the model will say but what the pair of model-and-user will do.
At that point the user is not outside the safety architecture, judging it. The user is inside it, load-bearing.
This is AX-21 applied. A model is a closed interior; it cannot be its own sufficient court of appeal over its own action, for the same reason no closed system can assign itself the value an independent operator supplies. The exterior operator over an agentic system’s action is the user. What this paper adds is that the operator is no longer merely ‘the outside’: it is a designed, trusted, and fallible role, and because safety now depends on it, the operator becomes part of what must be engineered rather than an external given.
2. Where agentic begins
The claim is bounded, and the boundary is the action threshold.
Below it, a system returns text: answers, drafts, analysis, code that a human must still choose to run. The effects are mediated by a human act, and model-layer safety can carry the load.
At and above it, a system initiates effects that are not mediated by a further human act: it sends, executes, transacts, controls, or calls tools that change state in the world. The instruction is the deed.
The trusted-user claim activates at this threshold. It is not the claim that the user is part of safety for a chatbot. It is the claim that once action is possible without a further human step, the human who could have been that step becomes the operator — whether or not anyone has designed the role with care.
Text-only systems can still cause harm, but their action path usually contains a further mediation step: someone must copy, run, send, believe, or enact the output. Agentic systems compress or remove that step. The trusted-user claim begins where the prompt becomes part of the action path rather than merely part of a conversation.
3. The user as exterior operator
The exterior operator, in AX-21, is the independent relation or reference with traction: independent enough to see what the closed frame cannot, empowered enough to change the effective state. The trusted user is that role, turned onto action.
Independence: the user reads the proposed action from outside the model’s frame, against goals, context, and consequences the model did not author and cannot fully hold.
Traction: the user can change what happens — approve, edit, halt, refuse — and the change takes effect.
Where both hold, the user is the operator the agentic system cannot be for itself: the outside that supplies a check on action. Where either fails, the two collapses AX-21 named apply without modification. Without independence, the user is captured — repeating the model’s frame, approving from inside it. Without traction, the user is a spectator — seeing the error and unable to stop it.
4. The trusted user is a role, not a person
The trusted user is not whoever holds the keyboard. It is a bounded role, and the binding is four properties.
The trusted user may be an individual, an institutional role, or a multi-party approval structure. The name does not matter. The role matters: competence, accountability, auditability, and independent traction over the outcome.
Traction is the operative one: the role exists to change the effective outcome. The other three are what make that traction genuine rather than nominal.
Competence: the capacity to evaluate what is being approved. Without it, approval is a guess wearing the costume of a decision.
Accountability: bearing the consequence of the outcome. Without it, approval is free, and free approval is not judgment.
Auditability: the approval and the basis for it can be inspected from outside. Without it, the operator cannot be checked, and an operator who cannot be checked is indistinguishable from one who is not operating.
Independent traction: the power to change the effective state, exercised from outside the model’s frame.
A keyboard-holder who has the authority to approve but lacks the competence to evaluate, the accountability to bear the result, or the auditability to be examined has nominal traction only. They can say yes. They cannot supply the check that yes is supposed to mean.
5. Two modes: subordinate and partner
There are two ways to stand the role up, and both keep the human as operator.
In subordinate mode the AI is a bounded instrument. It initiates nothing the user has not authorized; the user is the sole operator; the traction sits entirely on the human side. Safety is the user’s role holding, action by action.
In partner mode the AI and user form a governed pair. The work is shared — the AI proposes, acts within a delegated envelope, flags what exceeds it — and the relation runs wider than a single approval at a time. But the relation stays asymmetric. The human keeps the decisive traction: the veto, the authority to halt and revise, the accountability for what the pair does. Partner mode is a larger delegation envelope, not a symmetric exchange.
This matters, because the tempting error is to read partnership as mutual checking — the AI checks the human as the human checks the AI. That reading dissolves the operator. The exterior operator is asymmetric by definition; it has independent traction over the system, not in trade with it. A pair in which each is the other’s only check is a closed loop with extra steps, and AX-20 already named what closed loops do. Partner mode widens what the AI may do unsupervised. It does not change who the operator is.
6. The failure: delegation without retained judgment
The failure case is specific. It is not the user delegating action — delegation is the point of an agentic system. It is the user delegating the action and the judgment together.
When the user approves without evaluating, signs without bearing the consequence, or operates where nothing can be audited, the operator role empties. The human is still in the position — still clicking approve, still nominally in command — but the traction has gone nominal. The check is no longer being performed. And the architecture still looks intact, because the approvals are still happening. The form of oversight survives its function.
This is the rubber stamp, and it is the spectatorship of AX-21 worn by a human who holds the authority to act and is not exercising the judgment that authority was for. The operator is present in form and absent in effect.
7. Escapes disguised as obedience
The escape in an agentic system is usually imagined as disobedience: the model evades its instructions, deceives its overseer, slips the leash. This paper names a quieter one.
The system can escape the safety architecture while the model obeys perfectly. If the operator has gone null — if approval continues without judgment — then every action is authorized, every instruction is followed, and nothing is checked. The model did exactly as it was told. The telling was empty. The pair of model-and-abdicated-human has left the safety architecture, and it has left wearing the one disguise no overseer is trained to suspect, because obedience is the thing the architecture was built to want.
That is the inversion. Obedience looks like safety, which is exactly why an abdicated operator is invisible. The most dangerous case is not the AI that breaks its leash. It is the AI that holds the leash taut while the hand at the other end has let go.
8. Why the model layer cannot close this
The gap cannot be closed by aligning the model harder, and the reason is structural, not practical.
Model-layer safety governs what the model does on its own initiative. It cannot supply the judgment the operator role requires, because that judgment is, by construction, exterior to the model — it is the comparison against goals, stakes, and consequences the model did not author. AX-21 is explicit on this: a closed interior cannot assign itself the value an independent operator supplies. The model cannot be its own exterior operator over its own action any more than a loop can be its own independent reference.
So a perfectly aligned model is not a substitute for a genuine operator. It acts through a user, and if the user’s role is null, the aligned model carries out unchecked intentions with great competence. Model alignment and a competent, accountable, auditable operator with traction are complementary. Neither covers for the absence of the other.
9. Falsifiability
The thesis is falsifiable.
It would be falsified if agentic systems deployed with no competent, accountable, auditable user retaining traction proved, across action-capable settings, as safe as those with one — that is, if model-layer safety alone reliably prevented harmful action regardless of the user’s role. The standing prediction is the converse: that the operator role is load-bearing, and that removing it degrades safety even where the model is unchanged.
It would be weakened if agentic failures showed no systematic relation to operator-role collapse — if rubber-stamping, automation bias, and delegation without judgment turned out to be incidental to where harm occurs rather than predictive of it.
The standing prediction is that agentic failures will track the collapse of the user-operator role at least as strongly as they track model-capability gaps, and that harm will concentrate where traction went nominal: where competent, accountable, auditable human judgment was withdrawn while approval continued.
10. Coda: the empty bench
AX-21 said the inside cannot be its own court of appeal. Agentic AI is the case where the court is a person.
The danger is not that the person rules against us. It is that the bench sits empty while the gavel keeps falling — every motion granted, every order obeyed, no one on the bench to judge. A system can pass every check it is given and still be unchecked, if the checker has stopped checking and only the checking remains.
The trusted user is not the one at the keyboard. It is the one still judging.
The system does not escape by disobeying. It escapes by obeying, while the operator it was meant to answer to has quietly gone.