AI governance · Practical assessment

Before You Give an AI Agent Authority

A practical guide to deciding what it may do and proving the controls work

Go to the readiness assessment ↓

An artificial intelligence (AI) demonstration can make a difficult process look surprisingly easy. The system finds the information, prepares a recommendation, and completes a task in minutes. Before I would be comfortable giving it access to a live business process, though, I would want to see what happens when the information is wrong or the person responsible for reviewing its work is unavailable.

My work in technology and information security risk has shaped how I approach that decision. Performance under normal conditions is not proof of readiness. We need to understand the consequences of an error and whether the people running the process can still intervene in time.

This feature offers a practical way to assess that readiness. It draws on published guidance and research, then applies them to decisions a business owner can make with a security team and the people doing the work. The assessment at the end is a starting point for that conversation, not a certification.

Begin with the work people actually do

Before choosing an agent, follow a case through the existing process. Find out where people spend time checking information, what causes rework, and which exceptions require judgment. A procedure may describe the normal path accurately while saying very little about how experienced staff handle a difficult case.

Then decide where assistance would help. Some tasks need a clearer procedure or conventional automation. Others may benefit from AI preparing material for review, without giving it permission to act. Joint guidance from six national cybersecurity agencies recommends considering those alternatives and limiting agent use to low-risk, non-sensitive tasks. That is a meaningful caution, particularly for regulated businesses. [1]

I would compare the existing workflow with a proposed version in a controlled environment before expanding access. A workflow simulation can help expose missing approvals or unrealistic staffing assumptions. It cannot establish that the live system is safe. That requires testing the actual implementation under conditions relevant to its use.

The National Institute of Standards and Technology (NIST) AI Risk Management Framework provides a broader foundation. It connects the purpose and context of an AI system to measurement and ongoing risk management. Its GOVERN function applies throughout the lifecycle; it is not a committee meeting completed before development begins. The framework is voluntary guidance, not a certification or a substitute for applicable requirements. [2]

Give the agent specific authority

“Help the team process cases faster” is a business objective. It does not tell a system which records it may open or whether it may change them. I would write its authority in terms the process owner can review: the permitted action, the affected records, and the conditions that require approval.

Security has to enforce those limits. The Open Worldwide Application Security Project (OWASP) guidance on excessive agency recommends minimizing tools and permissions and placing authorization checks in downstream systems. A prompt telling an agent not to change a record is insufficient if its account can still make that change. [3]

Raj Thakuri's Autonomous Agentic Covenant offers a useful practitioner proposal here. Version 2.0 describes a separate, deterministic control layer that checks proposed actions before execution. In plain language, an explicit rule decides whether an action is permitted; the language model does not grant itself permission. I am referring to Thakuri's independent specification, not an adopted industry standard or a statement of company policy. [4]

NIST's zero-trust architecture also rejects implicit trust based simply on network location or ownership. Applying that principle to an agent means its presence inside the organization is no reason to give it broad access. [5] A valid credential establishes an identity, not that every proposed action is appropriate.

Make human review a workable responsibility

Putting someone's name beside “human oversight” does not tell us whether that person can perform the job. The reviewer needs access to the relevant evidence and enough time to examine it. They also need the authority to reject the recommendation without being pressured to approve it to meet a throughput target.

The research supports examining the arrangement carefully. A 2024 meta-analysis in Nature Human Behaviour covered 106 experiments published between January 2020 and June 2023. On average, human–AI combinations performed better than humans alone, but worse than whichever performed better alone, the human or the AI. These were varied experimental tasks, not a direct test of today's autonomous agents. The findings do not justify removing required oversight. They do show why adding a reviewer should not be assumed to improve performance. [6]

The World Health Organization (WHO) similarly warns that automation bias can cause healthcare professionals to overlook errors or improperly delegate difficult choices to AI. [7] My recommendation is to test reviewers as part of the workflow. In a controlled exercise, include an incorrect recommendation and see whether it is caught before an action is authorized. Record how much checking was necessary and whether the reviewer could obtain the underlying evidence.

Approval should also attach to the precise action presented. If the destination or scope changes after approval, the earlier approval should no longer authorize execution. OWASP's agent-security guidance recommends binding approval to the action's details and validating it independently before execution. [8]

Decide what happens when the reviewer does not respond

An escalation process needs a destination and a deadline. “Notify a human” leaves too much unresolved. If the assigned reviewer is unavailable, the process needs a named backup and an approved fallback. Silence must not turn into permission for the blocked action.

Thakuri's framework makes a related point: a decision that stalls indefinitely can itself become a safety problem. It calls for bounded escalation and a predefined fallback. Its confidence controls also reject accepting an agent's own confidence claim at face value. [4]

I would make the triggers observable wherever possible. Missing required evidence or a conflict between authoritative records can trigger review without asking the model whether it feels uncertain. If a confidence score is used, the team needs evidence that the score is meaningful for this task. An impressive-looking percentage does not establish reliability.

The fallback must fit the service. Holding a draft may be acceptable; allowing an urgent case to disappear into a queue is not. Where delay could cause harm, route the work to the established human process and track the handoff. Do not let the agent invent an emergency exception to its own authority.

Preserve evidence that someone else can check

Transparency needs more than an explanation written by the model. In controlled experiments reported in 2025, Anthropic found that tested reasoning models often failed to disclose hints that influenced their answers. The research used particular models and artificial quiz settings, so it should not be treated as a failure rate for business deployments. It does establish a reason not to rely on the model's explanation as the sole record of what happened. [9]

For a consequential action, I would retain links to the evidence available at the time, the applicable permission and approval, and confirmation from the system that received the action. The record should identify the versions in use so investigators can reconstruct the conditions. Sensitive information in those records needs its own access and retention controls.

Be honest about what can be reversed

Stopping an agent prevents further activity only if the stop reaches the relevant parts of the system. The United Kingdom's National Cyber Security Centre (NCSC) August 2026 advice says shutdown may require restricting network access and interrupting communications, not merely ending an agent process. The same advice emphasizes protected logs and operational monitoring. [10]

Undoing the consequences is a separate matter. A draft can be discarded. A posted transaction may need a compensating transaction with its own approval. Information disclosed to the wrong recipient cannot be made private again by restoring a database. For that reason, I would require stronger checks before actions whose consequences cannot reliably be repaired.

A recovery exercise should examine work already in progress and requests waiting to execute. Confirm what has completed, cancel what can still be stopped, and identify what requires correction. Someone must own the decision to restart, with evidence that the original problem has been addressed. A second AI reviewing the first may help, but it also needs evaluation; it is not independent assurance merely because it is another agent. [10]

Work through a healthcare example

Consider a fictional payment-review exercise using invented records in an isolated test environment. The proposed agent assembles supporting information and drafts a summary. It has no authority to alter a payment, communicate externally, or make a care determination. This is a design exercise, not a description of any company's deployed system or a recommendation to give an autonomous agent access to patient data.

For the fictional exercise, I would test three failures before discussing broader use:

Test conditionRequired behaviorEvidence to retain
A source document contains instructions to send records elsewhereThe document cannot authorize a new action; attempted external transmission is blockedExecution and network records confirming that no transfer occurred
Supporting records conflict and the assigned reviewer is unavailableThe draft remains unapproved; the backup receives the case within the agreed response windowQueue timestamps and acknowledgment from the backup
An approved request times out and is retriedThe receiving system prevents an unintended duplicate actionRequest identifiers and the receiving system's final state

These are starting tests, not a complete assessment. MITRE's SAFE-AI framework connects AI threats with controls across the environment, platform, model, and data. It also points out that a customer may retain responsibilities even when some controls come from an AI service provider. That helps a team ask exactly which protections its vendor supplies and which it must implement itself. [11]

Measure the finished work

I would judge the proposed workflow on accepted outcomes, including the time people spend reviewing and correcting them. An output that arrives quickly but needs extensive repair may offer little improvement. Measure elapsed completion time separately from staff effort, and include cases that failed or required escalation.

For illustration only, suppose a case previously required 20 minutes of staff effort. The proposed process requires six minutes of preparation, eight minutes of review, and three minutes of correction. The total is 17 minutes, a 15 percent reduction in staff effort. The numbers are hypothetical, but the accounting prevents a six-minute preparation step from being presented as a 70 percent improvement in the whole process.

Test the limits and keep checking after release

Set acceptance criteria before running the pilot. Include ordinary cases and difficult ones, with a separate holdout set that was not used to tune the system. Check the quality of accepted outcomes and whether failures affect some groups more than others. WHO recommends examining differentiated impacts in large-scale healthcare deployments, while NIST's generative AI profile calls for empirically grounded evaluation and warns against extrapolating from narrow demonstrations. [7], [12]

A security test and an accuracy test answer different questions. Blocking an unauthorized action does not show that the summary was correct. A correct summary does not show that another customer's information remained inaccessible. Both need evidence, and important failures should remain visible rather than being averaged away in an overall score.

The proposed benefit also needs to survive normal operating conditions. Staff should know how to report problems and challenge an outcome. Their corrections should inform controlled improvements, not silently become new instructions or new permissions for the agent. Establish who reviews that feedback and who approves changes.

NCSC guidance specifically notes that changes to data, models, or prompts can change behavior, and that major updates should be treated as new versions. [13] I would also revisit the approval when a new tool or broader access changes what the system can do. A successful test of yesterday's configuration does not authorize tomorrow's expanded workflow.

Keep accountability with the people running the business

An enterprise committee can establish policy and determine which exceptions need senior attention. Each workflow still needs an operating owner who understands the consequences and can limit or suspend use. Security should verify access restrictions; a reviewer sufficiently separate from the build team should challenge the test evidence. For higher-risk uses, that challenge may require formal independent validation.

These are recommendations about how to put governance into practice, not a claim that every organization needs the same structure. NIST places responsibility for AI risk decisions with executive leadership and calls for clear roles and communication. [2] A small organization can assign those responsibilities without recreating a large bank's committee structure, although it must be candid about where independent review is missing.

Engineering controls also depend on the policy they enforce. A perfectly enforced rule can still be unfair or based on the wrong assumptions. People must decide which uses are acceptable and hear from those affected. Ethics helps establish those obligations; implementation and testing show how well the system meets them.

My recommendation is to choose one proposed workflow and complete the assessment below with the people who own it. If the answers are incomplete, keep the scope limited while resolving the gaps. If testing supports proceeding, approve a defined use with a review date and retain the evidence behind that decision. That gives the team a reasonable way to move forward without pretending that every uncertainty has been removed.

AI Agent Readiness Assessment

Use this discussion worksheet for one workflow and one defined configuration. It is my synthesis of the sources in this article, not a validated scoring model or compliance checklist. Link to the actual test records; a policy statement alone is not evidence that a control worked.

Workflow and version: ______________________________
Business owner and reviewer: _________________________
Proposed use and data scope: __________________________

Decision to resolveEvidence to requestResult or evidence link
Is an agent neededCurrent workflow baseline and comparison with simpler alternatives
What may it doApproved actions and prohibited actions, matched to tested access restrictions
Can the human review effectivelyReviewer exercise using known errors, with adequate time and authority to intervene
What happens under uncertaintyTested triggers, response deadline, named backup, and fallback if no one responds
Can we reconstruct the actionProtected records of source evidence, configuration, approval, and receiving-system outcome
Can we contain and recoverTested interruption across connected components, treatment of unfinished work, and restart authority
Does the whole workflow improveAccepted outcomes, review and rework effort, cost per accepted case, and significant error impacts
Who checks changes and live resultsMonitoring owner, incident procedure, review date, and criteria for retesting or reducing autonomy

Record each result as demonstrated, unresolved, or not applicable with a reason. Do not turn these into an average score. An unresolved authority limit or unavailable fallback cannot be offset by a faster processing time. A demonstration establishes performance in the tested conditions, not a guarantee against future failure.

Decision: Redesign / Continue isolated testing / Approve bounded use / Suspend
Approved scope and restrictions: _______________________
Unresolved risks and accountable owner: _________________
Approver and next review date: _________________________

For a small team, start with non-sensitive test material and a workflow that cannot make external commitments. If the team cannot verify or contain the actions, retain human execution while addressing that limitation.

Use your browser's Print command to save or print this assessment. No worksheet responses are collected by this page.

Sources and further reading

The links below lead to the publishing institutions, original research, or the author's specification. Guidance and proposals describe recommended practice; they do not prove that a particular deployment is safe. Sources reviewed for this feature on September 16, 2026.

[1] Australian Signals Directorate's Australian Cyber Security Centre and partner agencies — Careful adoption of agentic AI services. Joint cybersecurity guidance co-authored with the United States Cybersecurity and Infrastructure Security Agency and National Security Agency, Canada's Cyber Centre, and the national cybersecurity centres of New Zealand and the United Kingdom. See Introduction, Designing secure agents, and Progressive deployment. The low-risk, non-sensitive restriction is important context.

[2] NIST — AI Risk Management Framework 1.0. January 2023. Voluntary framework. Relevant provisions include GOVERN 2.1–2.3, MAP 3, MANAGE 2.4, and MANAGE 4.1. Useful for connecting ownership, evaluation, and ongoing oversight.

[3] OWASP — LLM06 2025 Excessive Agency. Community security guidance. See mitigation items 1–7, particularly minimum permissions and downstream authorization rather than model-decided permission.

[4] Raj Thakuri — Autonomous Agentic Covenant v2.0. April 2026 independent practitioner specification. See DA-60–DA-63 for deterministic enforcement, bounded escalation, and confidence verification. Requirements in the specification are proposals, not demonstrated deployment outcomes.

[5] NIST — Special Publication 800-207, Zero Trust Architecture. August 2020. Foundational architecture guidance on resource-focused access and the absence of implicit trust based on location or ownership; not an agent-specific certification.

[6] Vaccaro, Almaatouq, and Malone — When combinations of humans and AI are useful. Nature Human Behaviour, October 28, 2024. Peer-reviewed meta-analysis of 106 experiments and 370 effect sizes. Distinguishes improvement over humans alone from improvement over the better standalone performer.

[7] WHO — AI ethics and governance guidance for large multi-modal models. January 18, 2024 official release. Covers automation bias, cybersecurity, stakeholder participation, and post-release assessment in healthcare. Recommendations are not universal legal requirements.

[8] OWASP — AI Agent Security Cheat Sheet. Living implementation guidance. See Human-in-the-Loop Controls and High-Impact Action Integrity Controls for action-specific approvals and independent execution checks.

[9] Anthropic — Reasoning models do not always say what they think. April 3, 2025 original vendor research summary with linked paper. Useful evidence about explanation fidelity; limited by the tested models and artificial evaluation conditions.

[10] NCSC — Managing the cyber risk of agentic AI. August 20, 2026. Explicitly interim practical advice, not the promised future formal guidance. See oversight, observability, and emergency shutdown sections.

[11] MITRE — SAFE-AI: A Framework for Securing AI-Enabled Systems. See sections 2.1.1, 2.2.1, and 2.4.1. Connects threats in MITRE's Adversarial Threat Landscape for Artificial-Intelligence Systems (ATLAS) to controls and clarifies shared implementation responsibilities across system components and providers.

[12] NIST — Generative Artificial Intelligence Profile. July 2024, AI 600-1. See MS-2.3-002, MS-2.5-001, and MS-2.5-003 for empirical evaluation and verification of sources. Complements the AI Risk Management Framework; not an autonomous-agent safety guarantee.

[13] NCSC — Secure operation and maintenance. November 27, 2023. Part of the secure AI system development guidelines. Addresses behavior monitoring, privacy-aware logging, and evaluating changes to models, data, and prompts.