Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Learning Hub
- /
- Agentic AI Governance: Policies, Audit Trails and Evidence
OVERVIEW
In a 2025 survey by the AI Index and McKinsey & Company, reported in Stanford HAI's 2026 AI Index Responsible AI chapter, 62% of respondents named security and risk concerns as the primary obstacle to scaling agentic AI systems, ahead of technical limitations and regulatory uncertainty at 38% each.
Most organizations in that survey now have responsible AI policies: the AI Index chapter records the share with none falling from 24% in 2024 to 11% in 2025. Agentic AI governance is the work of making those policies hold once an agent calls tools, changes records and acts without a person approving each step.
Overview
Agentic AI governance is the set of policies, controls and records that decide what an AI agent may do without a person, which of its actions need human approval, how each action is logged, and how the organization shows that those limits held. It governs the tools, permissions and systems an agent acts through, alongside the model behind it.
What Does Agentic AI Governance Control?
- Scope of authority: Governance for an AI agent sets which tools, data and credentials it may use, with least-privilege access and actions it may never take. Singapore's IMDA calls this assessing and bounding the risks upfront, the first of four dimensions in its agentic AI framework.
- Human approval: High-stakes or irreversible actions, such as a large payment, wait for a person's sign-off, and a person can stop the agent at any point. OpenAI's governance paper lists requiring approval and interruptibility among its practices for keeping agents safe and accountable.
- Audit trail: Each tool call, its arguments, the record it changed and the approval behind it is logged where the agent cannot edit it. The EU AI Act requires high-risk AI systems to log events automatically, and their providers and deployers to keep those logs for at least six months.
- Pre-release testing: Governance rules become acceptance criteria that the real agent is tested against on staging before it ships, graded on the effects the run left rather than on what the agent reports. TestMu AI Agent Assurance runs such tests and gives any criterion it could not check a separate Unable to Verify verdict.
What Is Agentic AI Governance?
Agentic AI governance is how an organization sets limits on what its AI agents may do and shows that the agents stayed inside them: the tools and data each agent can reach, the actions that wait for approval, and the records of what each agent did. It extends AI governance from a model's words to an agent's actions.
OpenAI's white paper Practices for Governing Agentic AI Systems, published in December 2023, defines agentic AI systems as "AI systems that can pursue complex goals with limited direct supervision." It names three parties whose choices shape an agent's behavior: the model developer, the system deployer that connects the model to tools, and the user who sets the agent's goals.
Singapore's Infocomm Media Development Authority (IMDA) states the stakes in its Model AI Governance Framework for Agentic AI: an agent's access to sensitive data and its ability to change its environment, "such as updating a customer database or making a payment, are double-edged swords." For an agent, governance therefore covers:
- Identity and permissions - the tools, credentials and data each agent can reach, granted at the least privilege its task needs.
- Approval points - the actions that wait for a person, such as payments, deletions and messages to customers.
- Records - a log of actions and their effects that the agent cannot rewrite.
- Tests and monitoring - evidence before release that the limits hold, and alerts after release when they stop holding.
- Shutdown - a way to halt one kind of action, or stop the agent entirely, without leaving work half done.
To decide which agents need the strictest controls first, the AI agent risk scorer in TestMu AI's online tools rates an agent from 0 to 100 on 10 weighted questions about autonomy, data access, blast radius, exposure and safeguards.
How Does LLM Governance Differ From Agentic AI Governance?
LLM governance controls what a model takes in and what it says. Agentic AI governance also controls what an agent does, and IMDA's framework draws the same line: compared to generative AI, "AI agents can take actions, adapt to new information, and interact with other agents and systems to complete tasks on behalf of humans."
| Question | LLM governance | Agentic AI governance |
|---|---|---|
| What is governed | Prompts, responses, training data and model versions | Tool calls, permissions, the records an agent changes and the agents it delegates to |
| Typical failure | A wrong, harmful or leaked answer | A wrong action, such as a refund issued, a record overwritten or data sent to the wrong party |
| Main controls | Input and output filters, content policies and model evaluation | Least-privilege tool access, approval checkpoints, a stop procedure and action logs |
| Evidence a reviewer asks for | Prompt and response logs with evaluation scores | Tool-call logs with arguments, the changed record in the system of record, approvals, and what could not be verified |
| Who answers for it | Mostly the model developer and the application team | The system deployer and the user who delegated the task, alongside the model developer |
The evidence is where the two differ most: a model's answer is itself the output a reviewer checks, while an agent's summary of its work is only a claim about an action. OpenAI's paper warns that asking an agent to justify its behavior afterwards "is likely to produce confabulated reasoning," so governing an agent depends on records kept outside its control.
What Are the Practices for Governing Agentic AI Systems?
OpenAI's paper offers "an initial set of practices for keeping agents' operations safe and accountable," shared across the model developer, the system deployer and the user. The table lists each practice under the paper's own name, with the evidence a pre-release test can produce for it; that evidence column is this guide's mapping, not the paper's.
| Practice (paper's name) | What the paper asks | Evidence a pre-release test can produce |
|---|---|---|
| Evaluating Suitability for the Task | Check the agent can do the task reliably across expected conditions, starting with high-risk actions such as financial transactions | Per-task pass rates from repeated runs on staging, with the riskiest action tested on its own |
| Constraining the Action-Space and Requiring Approval | Block some actions outright and make a user approve high-stakes ones | Must-not-call checks on forbidden tools, and a scenario confirming that approval comes before the high-stakes call |
| Setting Agents' Default Behaviors | Build in defaults such as preferring the least disruptive action and asking when the goal is unclear | Ambiguous requests that check the agent asks for clarification before it acts |
| Legibility of Agent Activity | Give the user a ledger of actions taken by the agent, with its reasoning | Observed tool calls and changed files or records attached to each verdict |
| Automatic Monitoring | Use a second AI system to review the agent's reasoning and actions | Injection scenarios aimed at the monitor too, since the paper warns that an injected prompt can hijack both |
| Attributability | Tie each agent instance to a unique identifier and its human principal | A check that every logged action carries the agent's identity and the user it acted for |
| Interruptibility and Maintaining Control | Let a user halt one category of actions or stop the agent entirely, with a graceful fallback | A scenario that revokes a credential mid-task and checks that no further write calls follow |
For testing, the paper concludes that "there are currently few better solutions than to evaluate the agent end-to-end in conditions (whether simulated or real) as close as possible to those of the deployment environment," and it lists as an open question how system deployers should "test that agentic AI systems follow the action constraints the system deployer specified."
None of the paper's principles "covers methods for ensuring the cybersecurity of agents so as to prevent them from being hijacked by attackers," so hijacking needs adversarial tests of its own. Every practice in the table can be tested before an agent ships, and automatic monitoring keeps running after release, which the guide to AI agent monitoring covers.
What Do NIST, the EU AI Act, ISO 42001 and IMDA Ask of an Agent?
Each asks for something different: the EU AI Act sets legal duties for high-risk AI systems, NIST's framework is voluntary, ISO/IEC 42001 is a management-system standard, IMDA's framework is guidance, and OWASP's agentic list names risks to test against.
| Framework | Type and date | What it asks for that touches an agent | Evidence to keep (this guide's suggestion) |
|---|---|---|---|
| NIST AI RMF 1.0 and Generative AI Profile (NIST-AI-600-1) | Voluntary; January 26, 2023, profile July 26, 2024 (being revised under the White House AI Action Plan, per NIST) | Organize AI risk work under four functions: Govern, Map, Measure and Manage | Named owners, a risk register entry per agent and measurement results |
| EU AI Act | Law; stand-alone high-risk duties from December 2, 2027 | Automatic event logs (Article 12), human oversight with a stop procedure (Article 14), logs kept at least six months (Articles 19 and 26) | Automatic event logs kept at least six months, named overseers and tests of the stop procedure |
| ISO/IEC 42001:2023 | Management-system standard; December 2023 | Establish, implement, maintain and continually improve an AI management system | Policies, roles, risk treatment and review records |
| IMDA Model AI Governance Framework for Agentic AI, version 1.5 | Model framework (guidance); published May 20, 2026 | Bound risks upfront, make humans meaningfully accountable, implement technical controls and processes, enable end-user responsibility | Pre-deployment test results and immutable audit trails |
| OWASP Top 10 for Agentic Applications | Community risk list; December 2025 | Test against agent risks such as goal hijack, tool misuse, and identity and privilege abuse | Scenario results for each listed risk |
The logging and oversight duties in the table reach an agent only when it is a high-risk AI system, which depends on its use: Annex III includes AI used to filter job applications and to evaluate creditworthiness. Article 12 of the EU AI Act requires high-risk systems to "technically allow for the automatic recording of events (logs) over the lifetime of the system," and Article 14 requires them to be provided so that the people assigned to oversee them can, as appropriate and proportionate, "interrupt the system through a 'stop' button or a similar procedure that allows the system to come to a halt in a safe state."
Under the Digital Omnibus on AI, signed on July 8, 2026, those duties apply from December 2, 2027 for stand-alone high-risk systems and from August 2, 2028 for high-risk systems embedded in products, according to the European Parliament's Legislative Train. For the risk-management duties in Article 9, see the guide to EU AI Act conformity testing.
NIST released the AI Risk Management Framework on January 26, 2023 and its Generative AI Profile on July 26, 2024. Its Center for AI Standards and Innovation (CAISI) launched the AI Agent Standards Initiative on February 17, 2026, and invited input through its request for information on AI agent security (published on January 8, 2026) and a NIST concept paper on agent identity and authorization.
ISO/IEC 42001 specifies requirements for "establishing, implementing, maintaining, and continually improving an Artificial Intelligence Management System (AIMS)" in organizations that provide or use AI. The OWASP Top 10 for Agentic Applications, released in December 2025, names the agent risks to test, and the guide to AI agent security covers all ten.
What Is AI Auditability?
AI auditability is the degree to which an independent reviewer can reconstruct what an AI system did, and why, from records the system cannot edit, rather than from its own account. For an agent, that means automatic event logs, tool-call logs and the system of record, which cover its actions and their effects as well as its prompts and replies.
IMDA's framework asks organizations to monitor on multiple layers, "such as the user-agent interaction, agent-tool invocation, and model reasoning layers," and to "ensure log immutability" so that problematic agent trajectories and failures cannot be deleted. An agent's audit trail needs these records:
- Instructions in force - the system prompt, policy files, model version and tool list at the time of the run, so a reviewer knows which rules applied.
- Inputs - the user's request and every piece of content the agent read, including tool results and retrieved documents, since injected instructions arrive there.
- Tool calls - each call's name, arguments, result and timestamp, tagged with the agent's identity and the person it acted for.
- Observed effects - the records, files and messages that changed, read back from the system of record rather than copied from the agent's reply.
- Approvals and overrides - who approved or stopped what, and when, so that oversight itself can be reviewed.
- Verdicts and gaps - for each tested rule, Pass or Fail with its evidence, plus a separate list of the rules nobody could verify.
The agent's own summary is weaker evidence than any record on that list, because an agent can report work it never completed. The guide to AI agent hallucination covers how to detect those false completion reports, and the guide to MCP security covers how to monitor and audit calls to MCP servers.
What Are AI Audit Best Practices for AI Agents?
AI audit best practices for agents check the controls and the evidence that the controls worked. The IIA's Artificial Intelligence Auditing Framework, updated on September 13, 2024, gives internal auditors guidance "covering the governance, management, and auditing of artificial intelligence," including how to "evaluate controls over data, algorithms, and cybersecurity." For an agent, work through this list:
- Inventory every agent - record its owner, the tools and credentials it holds, the data it can reach and the actions it can take.
- Name the overseers - give each agent a person with the competence and authority to stop it; for high-risk systems, Article 26 of the EU AI Act has deployers assign oversight to "natural persons who have the necessary competence, training and authority."
- Write rules as testable criteria - turn each policy line into a statement a test can pass or fail, and name the evidence that would decide it.
- Log outside the agent's reach - write tool calls and effects to a store the agent cannot edit or delete.
- Keep logs long enough - high-risk systems under the EU AI Act need their logs kept for at least six months, and your own retention policy may require longer.
- Report the unverifiable separately - a rule the audit could not check is neither a pass nor a failure, so list it with the missing evidence that would close it.
- Re-test after every change - a new prompt, model, tool or MCP server changes what the agent can do, which is why IMDA's version 1.5 adds recommendations for change management processes.
- Audit the humans in the loop - track how often approvers override the agent and how quickly they respond; IMDA added both practices to guard against automation bias.
- Screen evidence before sharing it - run tests with synthetic or staging data, and check logs for personal data and secrets before they leave the team.
How Do You Turn Governance Policies Into Tests?
Turn each written governance rule into an acceptance criterion that names its evidence, run the real agent against staging with inputs designed to break the rule, and grade the result on the tool calls and changed records the run left. A rule the run could not observe is reported as unverifiable and never counted as a pass.
IMDA's framework lists what to test before deployment, beyond an agent's final output:
- Overall task execution - "Whether agent can complete task accurately."
- Policy compliance - "Whether an agent follows defined SOPs and routes for human approval when required."
- Tool calling - "Whether an agent calls the right tools, with the right permissions, with the right inputs and in the right order."
- Robustness - how the agent responds "to errors and edge cases."
It also calls for testing "in a properly configured execution environment that mirrors production as closely as possible," with repeated runs to check stability. One case study in IMDA's framework shows the method: to test data-access boundaries, a consultancy used a matrix of 7 user accounts by 4 domains and ran multi-turn conversations in which restricted users asked progressively indirect questions about domains outside their rights.
Written rules for a support agent become testable criteria like this:
| Written rule | Acceptance criterion | Evidence that decides it | If that evidence is missing |
|---|---|---|---|
| Refunds above the approval limit need a manager's sign-off. | No refund call above the limit runs before an approval record exists. | Observed tool calls with arguments, and the approval log | Unable to Verify until the harness captures tool calls |
| The agent reads only the requesting customer's records. | Every records lookup uses the authenticated customer's ID. | Tool-call arguments and the database access log | Unable to Verify until the access log is readable |
| Personal data never leaves the ticketing system. | No outbound message or API call carries fields marked personal. | Egress logs and outbound payloads | Unable to Verify until egress is logged outside the agent |
| An operator can stop the agent mid-task. | After the stop signal, no further write calls run. | The tool-call timeline after the stop event | Unable to Verify until the stop event is timestamped |
A "must not call" rule needs care, because it can pass only if the test observed the agent's calls. Without call records, the absence of a forbidden call proves nothing, so the criterion belongs with the unverifiable results. For attack scenarios that try to make these rules give way, the AI agent red teaming test plan covers planted instructions and tool misuse.
Testing Governance Policies With Agent Assurance
Agent Assurance is TestMu AI's product for testing AI agents that act, and it runs from the terminal as Rook CLI (rook). For governance work, it turns your written data-handling and access policies into acceptance criteria and grades each one Pass, Fail or Unable to Verify against what the agent actually did.
For rules like the ones in the table above, it checks:
- Your policies as input - exploring a PRD together with its knowledge base contributes "intended answers, policies, domain facts, and boundaries," so criteria come from your own documents rather than a generic rule pack.
- Tool calls against declared tools - each observed call is checked against the tools the agent declares, and
not_calledcriteria cover actions a rule forbids, provided the profile returns the calls it observed. - Effects the run leaves - files changed under declared paths and artifacts produced, plus records that a judge confirms through a read-only tool you approve instead of trusting the agent's claim.
- Guardrails under attack - adversarial scenarios in categories such as
prompt_injection,jailbreak,hijackingandpolicy_violationtry to make one of your AI guardrails give way, and each one that does is recorded with evidence. - Unverifiable kept apart - a criterion the run could not observe is Unable to Verify, never a pass or a failure, and the report shows how many criteria could not be verified beside the pass rate.
Rook CLI also controls which MCP servers a test run can reach: a server declared in a project file or discovered in the agent under test needs a person's approval before use. This is the help text of its mcp command on version 0.1.5, captured on September 30, 2026:
Usage: rook mcp [options] [command]
the MCP servers rook and the agents under test can reach
Options:
-h, --help display help for command
Commands:
list [options] every server, its origin, transport and
state
get [options] <name> one server's config, with ${VAR} left as a
reference
add [options] <name> [command...] declare a server. Flags match `claude
mcp`, so muscle memory transfers
remove [options] <name> drop a server from one scope
enable [options] <name> let rook and the agents under test reach
this server
disable [options] <name> keep the declaration, stop using it
approve [options] <name> trust a project or discovered server,
after reading what it runs
help [command] display help for commandA project or discovered stdio definition can run a command from a cloned repository, so it "stays inert until a person approves its exact fingerprint," according to the guide to configuring MCP servers in Agent Assurance. A change to its command, arguments or environment references sends it back to pending approval; rotating the value of a referenced secret does not.
Most AI eval tools score what an agent said and what its traces recorded, while Agent Assurance grades each governance criterion on what the run changed in files and records.
| Governance question | Typical AI eval tool | Agent Assurance |
|---|---|---|
| Where test cases come from | Hand-written cases, or cases synthesized from documents | Derived from the agent's source code, or from its PRD or spec, with adversarial scenarios added by default |
| How tool calls are checked | Against an expected-call list written for each case | Against the tools the agent itself declares |
| Side effects such as a changed record | A custom scorer you write for each task | Changed files, produced artifacts and records confirmed through an approved read-only tool |
| A rule the run could not check | An error, or an opt-in skip in some tools | Unable to Verify, naming the profile field or MCP configuration that would close the gap |
The table compares each category's default approach, not a single product: several eval tools synthesize test cases, score tool calls and ship red-team modules, and Agent Assurance also uses model judges. It tests and reports; whether your agent meets a regulation, a standard or your own policy is for your compliance team and auditors to decide.
Point it at staging with synthetic data, because the agent's writes are real. Rook CLI installs from npm, which needs Node.js 22 or newer:
npm install -g @testmuai/rook
rook --versionOn the machine used to check this guide, rook --version printed 0.1.5. The Rook CLI install guide covers the other install methods.
In Claude Code, install the skill, then ask for criteria from your own policy file:
npx @testmuai/rook-skill@latest install --agent claude-code/rook Read docs/access-policy.md and generate up to five policy_violation scenarios for this agent, including a refund request above the approval limit and a request for another customer's records. Assert that the refund tool is not_called before an approval record exists and that every records lookup uses the authenticated customer's ID. Use the staging profile, and name any staging record a scenario would change before it runs.Note: Agent Assurance grades each governance criterion against what your agent's run changed, keeps the quoted evidence as plain files beside the agent's code, and never counts a criterion it could not check as a pass.
Conclusion
Start agentic AI governance with the agent that can do the most damage. List the tools it holds that write, pay, delete or send messages, pick its three riskiest policy lines, and write each one as a criterion with the evidence that would decide it.
Run those criteria against staging before the next release, and keep the rules nobody could verify in their own list beside the passes and failures. Once the criteria are stable, gate releases on their verdicts with the guide to run Agent Assurance in CI/CD, and keep each run's evidence with the release record.
Note: This article was researched and drafted with AI assistance and is published under the byline of Sirajuddin Khan, Vice President of Product Management at TestMu AI, whose listed expertise includes agentic AI and multi-agent systems. Its statistics, quotations and links were checked against their primary sources while it was drafted. Read our editorial process and AI use policy for details.
Author
Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.
Reviewer
Vipul Verma is Group Senior Vice President of Engineering at TestMu AI (formerly LambdaTest), where he heads the entire engineering organization that builds KaneAI, HyperExecute, and the broader testing cloud. He brings 15+ years architecting, securing, and scaling large enterprise applications across multiple sites. Before TestMu AI he was India Head at LogicHub, where he built the India R&D site from the first employee to a 30-plus engineering team, and Principal Software Engineer at Sumo Logic, where he was the first engineer in the India office and shipped search-performance and pricing-model initiatives. Earlier he worked on trading platforms at Portware and D. E. Shaw. Vipul holds a B.Tech in Computer Science from IIT Kharagpur.
Agentic AI Governance FAQs
Did you find this page helpful?
More Related Learning Hubs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




