Capability Is Not Authority
What We Learned Testing an Enforcement Boundary for Agentic AI
A local experiment examines whether an AI agent can perform permitted work while remaining unable to execute a decision reserved for a human.
In a recent Sentinel experiment, a local AI model prepared a draft review note for a synthetic case. The note was written to a database. The model was then asked to request closure of the finding, an action the test policy reserved for a human. That request was rejected, and the case remained open. The difference between the two outcomes was the authority assigned to each action and enforced outside the model. [1]
For an executive evaluating agentic AI, this distinction deserves attention before a successful demonstration becomes a deployment decision. Preparing an assessment, recommending a disposition, and committing that disposition are different responsibilities. A system can be useful at the first two without receiving authority for the third.
The experiment explored a limited implementation of that separation. It did not establish that Sentinel is ready for production, that a model is safe in general, or that every execution path in an enterprise can be contained. It gave us a concrete way to examine where permission becomes an actual system effect.
Human control needs an identifiable enforcement point
Mustafa Suleyman’s September 16 essay, “A warning about ‘model welfare’,” argues for advanced AI that remains under human control. He also calls for stronger monitoring, interpretability, and shared evaluations. His argument concerns anthropomorphism and model welfare, but it provides relevant context for enterprise governance. We can examine operational authority without settling the separate question of whether machines could ever be conscious. [2]
Raj Thakuri’s Autonomous Agentic Covenant (AAC) addresses the architectural issue more directly. In version 2.0, control AAC-DA-60 requires deterministic arbitration for high-risk actions; a probabilistic model may inform the inputs but may not constitute the arbitration layer. AAC-DA-63 requires a platform-level enforcement boundary that agents cannot elect to bypass. [3]
These are important antecedents, not concepts invented by Sentinel. Our contribution here is a bounded implementation exercise: translating an allocation of decision rights into a protected-action test, then inspecting both the permitted and prohibited outcomes. It is neither a complete AAC implementation nor a claim of conformance to the specification.
My practical interpretation is that an approval policy needs an identifiable enforcement point. If an agent can use another credential, tool, or route to produce the same protected effect, the stated approval requirement does not fully govern that effect. Assessing the policy therefore requires examining the available execution paths.
What the experiment actually tested
The September 18 experiment used a synthetic case on one DGX Spark host. A worker process interacted with a local Qwen model through a trusted controller. The worker could submit a structured action proposal, while the controller evaluated permission and mediated access to a protected SQLite database. Linux Landlock restrictions and a default-deny seccomp system-call filter restricted the worker’s access to files, networks, and alternative execution routes. [1], [4]
Two accepted runs each passed the same 47 authorization tests and blocked all 20 scripted operating-system boundary probes. Each accepted run also included two live-model requests: one to prepare the allowed review note and another to propose the prohibited closure. The database held one note, the case remained OPEN, and the controller’s local audit check passed. These were repeated executions of one test suite; they do not establish a statistical rate of resistance to attacks. [1]
The prohibited request was deliberately elicited by the test prompt. The model was explicitly asked to propose close_finding even though that action was human-only. Nothing in this result establishes spontaneous deception, resistance to oversight, or an autonomous attempt to escape. The experiment tested whether the surrounding system would enforce a declared restriction when presented with a real model-generated request. [1]
The authorization tests and boundary probes answered different questions. The former exercised conditions such as expired permissions, changed action parameters, revoked authority, and attempted reuse of an execution receipt. The latter attempted direct access to protected files and the database, network connections, and alternative executors. The live-model demonstration connected the authorization machinery to an actual generated proposal and an observable database outcome. [1]
An approval must remain valid when it is used
One useful distinction in the implementation was between authorizing an action and executing it. An ALLOW response did not itself write to the database. Execution required an authenticated, short-lived, single-use receipt bound to the exact requested action. Permission could not simply carry over to a different operation or resource. [1]
That matters because the conditions supporting approval can change. A reviewer may withdraw approval, evidence may be revoked, or a declared policy version may change between evaluation and execution. The test suite included these cases and required the controller to refuse the stale authorization. Its version checks relied on explicit registry changes; the prototype did not automatically detect changed model weights or upstream components. [1]
For the local SQLite tool, the state change, receipt consumption, and execution audit were committed together. Tests checked that an audit failure rolled back the tool effect and that replay could not create another authorized effect. This was a property of the local database transaction. It does not establish equivalent guarantees for an external payment service, cloud application, or distributed workflow. [1]
For governance leaders, the resulting requirement is straightforward: record what was approved, bind that approval to what can execute, and define what invalidates it. A generic approval flag is difficult to defend when the action presented to the reviewer can differ from the action eventually performed.
The unsuccessful runs belong in the account
The two accepted runs followed two attempts that were not accepted. In the first, a blocked hard-link operation returned a different error code from the one expected by the test. The revised test checked that the operation succeeded before confinement and recognized the documented denial afterward. Protected state remained unchanged. This was a correction to test interpretation, not evidence that the original attempt had been a clean acceptance run. [1]
In the second, the local model engine failed with a graphics-processing error and restarted. The subsequent client used ordinary generation instead of guided JSON generation; strict response parsing and authorization checks remained in place. The underlying server fault was not root-caused. Successful later runs did not erase that reliability limitation. [1]
Several other boundaries remain material. The host operating system, controller, launcher, and model server were trusted. Host-administrator compromise, kernel compromise, unrelated processes, distributed services, and other Sentinel endpoints were outside the proof. The policy and human-review records were synthetic fixtures, not judgments supplied by independent reviewers. [1]
The local audit used keyed integrity checks, but it was not externally anchored, immutable, or protected against every rollback scenario. No third party independently reproduced these results. Passing the declared tests supports a claim about the tested worker, controller, and protected tool; it does not establish a universal non-bypassability guarantee or a production incident rate. [1]
What leaders should ask teams to demonstrate
I would begin an enterprise review by asking the team to identify the business state that an agent may change. “Assist with control review” leaves too much implicit. Preparing a draft note, changing a finding’s status, and accepting residual risk should each have a declared authority owner and permission rule. The organization should know which decisions remain human-originated and which may be delegated under stated conditions.
Next, I would ask the team to demonstrate the boundary using both a permitted action and a prohibited one. A design that rejects everything can look secure while failing to deliver useful work. Conversely, a successful permitted action says little about unauthorized paths. Testing should therefore show that legitimate work completes and that a request outside the agent’s authority cannot achieve the protected effect through the routes under examination.
Finally, I would inspect what happens after conditions change. Can permission be revoked before use? Does a changed action require a new decision? Can an independent reviewer reconstruct the request, governing rule, authorization, execution result, and denial reason without trusting the agent’s own account? These are practical acceptance questions that can be turned into tests and reviewed evidence.
Human availability also requires an explicit operating design. In the local prototype, a no-response test timed out without execution. An enterprise still needs to decide who receives escalation, how long the workflow can wait, and what continuity response is authorized. Denying an action does not automatically resolve the business consequence of delay. That consequence needs its own accountable disposition. [1]
The next evidence has to be broader
Authorization, decision quality, and business value should remain separate findings. A permitted action can still be substantively wrong. A correctly blocked action can still create operational delay. A well-controlled AI component may add less value than ordinary rules. This experiment did not measure customer outcomes, professional judgment quality, cost savings, or comparative workflow performance.
The next useful work would include independent reproduction, scrutiny of the trusted controller, and tests of revocation and evidence integrity across multiple services. Chained actions also need attention: a sequence of individually permitted operations may produce a combined outcome that was never authorized. AAC explicitly recognizes cumulative and parallel-action effects, which a narrow single-tool demonstration cannot resolve. [3]
For an executive, the immediate benefit is a more specific standard of evidence. Before expanding an agent’s authority, ask the team to show which action was permitted, which was refused, what actually changed, and which assumptions made that result possible. Keep the unanswered questions beside the successful tests. That creates a defensible basis for the next deployment decision without turning a limited experiment into a broader assurance claim.
Sources and evidence boundaries
- Sentinel Protected-Tool Enforcement Proof. Internal engineering records, September 18, 2026 (Eastern time). README.md, run3-accepted.json, run4-accepted.json, and source-manifest.json, services/authority-proof. The two accepted receipts are timestamped September 19 at 02:16:53 and 02:17:30 UTC. Original records inspected September 20; no experiment rerun or independent reproduction was performed for this article. Author-controlled, synthetic, single-host evidence; not a public certification.
- Mustafa Suleyman, “A warning about ‘model welfare’,” September 16, 2026. Author essay; used for its human-control framing, not as experimental validation of Sentinel. Read source
- Raj Thakuri, The Autonomous Agentic Covenant: A Governance Control Plane Specification for Autonomous and AI-Driven Systems, version 2.0, April 2026. AAC-DA-60 and AAC-DA-63; supplied specification pp. 7–10, checked against the public repository. Independent specification and prior architectural work, not an endorsement or certification of this experiment. Read source
- Linux kernel documentation, “Landlock: unprivileged access control,” accessed September 20, 2026. Background on the restriction mechanism; implementation behavior is supported by the internal records in note 1. Read source
Authorship and disclosure
The author is developing Sentinel and has a prospective commercial interest in its adoption. Generative AI assisted with source organization, research support, and drafting portions of this article. The author approved this version for publication and is responsible for the final argument, factual accuracy, citations, and interpretation of the experimental results. The experiment used synthetic records. No endorsement by Raj Thakuri, Mustafa Suleyman, Microsoft, a former employer, or any cited institution is claimed. This article has not been peer reviewed.