How are AI agents audited when they take real-world actions?
Auditing AI agents is moving from access control to checking each action, but the field still offers frameworks and prototypes rather than proven audit practice.
Covers: The methods, standards and institutions used to audit AI agents that act in the world — such as tool use, transactions, robotics and autonomous workflows — including logging, traceability, third-party evaluation and regulatory frameworks. It does not cover general AI safety theory or model training audits that don't involve real-world actions.
Also answers: How do AI agents get audited when they take real-world actions? · How do you audit an AI agent that acts in the real world? · What does auditing AI agents with real-world actions involve? · Who audits AI agents that take actions?
- One page for this question6 other ways of asking lead here
- 7 independent sourcesEvery claim links to what supports it
- 3 connected pages3 changed this week
- Clean discussionScreened before anything appears
The short answer
Interpretation AI-prepared starting mapAuditing AI agents that act in the world is being approached from several directions at once: extending security maturity models from access control to action-level trust evaluation, building sandboxed environments with auditable identity layers that log every agent action, and governance-first architectures that gate actions before they execute. In healthcare, the emphasis is on predeployment accountability charters, bounded autonomy, postdeployment surveillance and named accountable parties. The evidence base is mostly architecture proposals, expert reviews and simulations rather than independent audits of deployed agents, so the field currently offers frameworks and prototypes more than proven audit practice.1234
- Evidence 24
- Interpretation 6
Did this answer your question?
Be the first to voteIn brief
Auditing is shifting from access control to action-level evaluation: AI-ZTMM defines 41 security Functions and seven Action Risk Factors, and covered 75.4% of the gap where the CISA ZTMM had no directly relevant controls (59.7% of analysed attack stages).1
Evidence-backedVerifiable identity and per-action logging are a common substrate: Astral uses OAuth 2.0 Token Exchange to create verifiable actor claims that log every agent action, validated in controlled simulations against Tool Poisoning Attacks.2
Evidence-backedGating actions before execution can be made planner-invariant: AEGIS admitted zero unsafe actions across 4,000 trajectories and four planner families, while a confidence-threshold baseline ranged from 0.03 to 0.998 false-allow depending on the planner.3
Evidence-backedIn healthcare, the proposed audit instruments are predeployment: an accountability charter naming intended and excluded use, validation evidence, oversight design, escalation pathways, auditability, subgroup performance review and accountable parties.4
Evidence-backedThe classic IT audit frame — evidence-based examination of management controls for asset safeguarding, data integrity and effective operation — remains the baseline that action-level agent auditing extends.5
Evidence-backed
At a glance
The picture in numbers
Live · updated just now
- Lowest planner0.03 false-allow rate
- Highest planner1 false-allow rate
- Mortality reduction0.8–12%
- Major adverse cardiovascular event reduction4–12%
The evidence behind it
7 sources- Reviews of many studies1
- Other studies and data5
- Background1
Published in 2026
| Source | Kind | Year |
|---|---|---|
| An Action-Centric Zero Trust Maturity Model for Agentic AI Environments. | Other studies and data | 2026 |
| A Secure Sandbox Environment for Orchestrating Medical AI Agents Using Model Context Protocols and Role-Based Access Control. | Other studies and data | 2026 |
| LATTICE: a governance-first architecture for authorized autonomous AI operations. | Other studies and data | 2026 |
| Medical AI Agents for Clinical Decision Support: Viewpoint Using the Planning, Action, Reflection, and Memory (PARM) Analytical Lens. | Other studies and data | 2026 |
| From assistance to autonomy: AI agent systems in cardiovascular medicine-a review of paradigms, architectures, and clinical translation. | Reviews of many studies | 2026 |
| Accountability for large language models in health care. | Other studies and data | 2026 |
| Information technology audit (Wikipedia) | Background | Unknown |
The community around it
No one has added to this page yet. Firsthand experience, a newer study or a different reading of the numbers would show up here, credited to you.
What it means for you
Which fits you?
Pick the situation closest to yours. Each answer says what it rests on.
If you are assessing an organisation's controls over what agents do after they have been granted access
an action-level maturity model such as AI-ZTMM gives you a structured self-assessment: 41 Functions, ten threat categories, 31 requirements, an Action Space and seven Action Risk Factors.1
Evidence-backedIf you need to reconstruct who or what performed each agent action
an auditable identity layer issuing verifiable actor claims per action, as in Astral's use of OAuth 2.0 Token Exchange, is the pattern the evidence supports — though it has so far been validated in controlled simulations.2
Evidence-backedIf you are deploying agents whose actions are hard to reverse and you cannot rely on the model's own confidence
a governance-first enforcement layer that gates actions before execution is the approach tested by LATTICE, which reported zero unsafe actions invariant to the planner; note that this was at a conservative operating point that auto-allowed no action.3
Evidence-backedIf you are introducing an agent into a clinical or other high-risk workflow
the proposed predeployment accountability charter — intended use, excluded use, validation evidence, oversight design, escalation pathways, auditability, subgroup performance review and named accountable parties — is the instrument to adopt, integrated into existing review, accreditation and procurement.4
Evidence-backedIf you are designing oversight for a medical agent
bounded autonomy, auditability, verification protocols, postdeployment surveillance and clear accountability structures are the mechanisms named, with evaluation covering end-to-end task reliability, escalation behaviour and performance under deployment shifts.6
Evidence-backedIf you work in a health system with limited regulatory infrastructure
the responsibility gap is noted as falling inequitably on low- and middle-income countries, which is the context in which earned delegation and predeployment charters are proposed as adoptable standards.4
Evidence-backedIf you are setting up an audit function for agentic systems from scratch
the conventional IT audit definition — evidence-based examination of management controls for safeguarding assets, data integrity and effective operation, often alongside financial or internal audit — is the baseline to build on.5
Evidence-backedThe full story · 5 chapters
01
Extending security maturity models to the actions agents take
AI summary:AI-ZTMM extends zero trust maturity models to action-level trust evaluation, scoring an organisation's own controls against fixed Functions and risk factors.
Evidence-backed: Existing Zero Trust Maturity Models such as the CISA ZTMM focus on how resources are accessed and give limited guidance on evaluating actions taken after access is granted. AI-ZTMM extends CISA's five-pillar structure to action-level trust evaluation, defining forty-one security Functions based on ten threat categories and thirty-one security requirements, and introducing an Action Space and seven Action Risk Factors for organisational self-assessment. Its scope covers software agents and the software action layer of agents in IoT, robotic and OT/ICS environments.1
Evidence-backed: The model was refined through reviews by eleven domain experts and evaluated against thirty-eight MITRE ATLAS case studies. The CISA ZTMM lacked directly relevant controls for 59.7% of the analysed attack stages, whereas AI-ZTMM addressed 75.4% of that gap, for a combined direct coverage of 84.4%. The authors present AI-ZTMM as complementing rather than replacing the CISA ZTMM.1
Interpretation: Read as an audit method, this is a self-assessment instrument: an organisation scores its own action-level controls against a fixed set of Functions and risk factors. That makes it useful for structuring internal review, but it is not an external attestation and the coverage figures come from mapping case studies rather than from auditing live agents.1
02
Sandboxes, verifiable identity and logging every action
AI summary:Astral sandboxes medical AI agents and uses verifiable identity claims to log every action, validated only in controlled simulations.
Evidence-backed: Astral is a secure sandbox environment for orchestrating medical AI agents using an Orchestrator-Specialist model, intended for controlled experimentation and evaluation in clinical settings. It enforces safety and accountability through an auditable identity layer using OAuth 2.0 Token Exchange, which creates verifiable actor claims to log every agent action. A generative visual layer built on persistent WebSockets and Model Context Protocol supports low-latency multimodal interaction.2
Evidence-backed: Controlled simulations validated that the architecture neutralises Tool Poisoning Attacks and improved transport performance. The authors frame this as a pragmatic foundation for trustworthy clinical AI, and note that current agentic systems are unsuitable for clinical use because of insufficient security mechanisms and the absence of visual components needed for effective use.2
Interpretation: The audit-relevant idea here is that identity and logging are the substrate: if each action carries a verifiable actor claim, then a later review can reconstruct who or what acted. The limitation is that the validation is simulation-based, so the logging layer has not been shown to hold up under real clinical load or adversarial conditions outside the sandbox.2
03
Governance-first enforcement: gating actions before they happen
AI summary:LATTICE gates agent actions before execution, admitting zero unsafe actions across planners, though at an operating point that auto-allowed nothing.
Evidence-backed: LATTICE is a governance-first architecture for authorised autonomous AI operations. In a pre-specified, planner-invariant safety evaluation across four frontier planner families (GPT-5, Claude Sonnet 4.6, Gemini, Grok-4; 4,000 trajectories), a confidence-threshold baseline's false-allow rate ranged from 0.03 to 0.998 across planners, while the AEGIS reference implementation admitted zero unsafe actions (false-allow 0.0, recall 1.0) invariant to the planner, at a conservative operating point that auto-allowed no action. A separate live run governed real operating-system actions with zero unsafe executions.3
Evidence-backed: Governance latency is reported as low and host-specific: on an Apple M4 Pro, policy evaluation at p50 ≈ 6.2 μs and full gated enforcement at p50 ≈ 0.7 ms including audit I/O. The authors position LATTICE as a pathway for responsible deployment in defence, critical infrastructure and regulated industries where authorisation requires verifiable governance rather than trust in AI behaviour.3
Interpretation: The wide spread in the baseline's false-allow rate across planners is the striking result: the same confidence threshold behaved very differently depending on which model was driving. That is an argument for auditing the enforcement layer rather than the model, but the zero-unsafe-action result was achieved at an operating point that auto-allowed nothing, so it says little about throughput or usefulness in a permissive setting.3
04
Clinical governance: charters, bounded autonomy and surveillance
AI summary:Clinical proposals emphasise predeployment accountability charters, bounded autonomy, postdeployment surveillance and named accountable parties.
Evidence-backed: A viewpoint on medical AI agents argues that agentic architectures incorporating planning, action, reflection and memory represent an evolution beyond rule-based, machine learning and multimodal clinical decision support. It identifies the governance mechanisms required for responsible implementation: bounded autonomy, auditability, verification protocols, postdeployment surveillance and clear accountability structures, while emphasising agentic AI as a supervised workflow support paradigm rather than autonomous modification of clinical judgment.6
Evidence-backed: Safe implementation is said to require technical safeguards, institutional governance, regulatory clarity, and evaluation approaches that assess end-to-end task reliability, escalation behaviour and performance under deployment shifts.6
Evidence-backed: A separate proposal introduces earned delegation as a predeployment standard: authority should be delegated only when evidence is proportionate to clinical risk, substantive human oversight is integrated and resourced, and responsibility for foreseeable failure modes is assigned in advance. To operationalise it, health systems should adopt a predeployment accountability charter specifying intended use, excluded use, validation evidence, oversight design, escalation pathways, auditability, subgroup performance review and named accountable parties, integrated into existing institutional review, accreditation and procurement structures.4
Evidence-backed: The same source notes that the responsibility gap distributes harm inequitably, particularly in low- and middle-income countries where regulatory infrastructure for digital health is still developing.4
Evidence-backed: In cardiovascular medicine, the C.A.R.D.I.O. framework (Clinical validation, Auditability, Risk stratification, Data privacy, Integration, Ongoing vigilance) and the CURACO framework (Clinical safety, Understanding, Research-informed care, Authentic patient-centred approaches, Conscientious ethics, Optimised technology) are offered as governance structures for responsible deployment. A rapid systematic review of 13 studies including 22,641 participants found that 85% of AI interventions improved cardiovascular outcomes, with mortality reductions of 0.8%–12% and major adverse cardiovascular event reductions of 4%–12%. The authors argue the transition from assistance to autonomy requires fit-for-purpose evaluation, transparent interpretability and robust governance, with agentic systems augmenting rather than replacing clinical judgment.7
Interpretation: These clinical frameworks converge on a similar audit shape: define permitted and excluded use in advance, require evidence proportional to risk, keep a human in the loop with a resourced escalation path, and name who is accountable. The outcome figures cited are for AI interventions broadly and predate agentic deployment, so they should not be read as evidence that audited agents improve outcomes.647
05
The older IT audit baseline
AI summary:Classic IT audit remains the baseline: evidence-based examination of management controls tied to attestation.
Evidence-backed: Information technology audit is an examination of the management controls within an IT infrastructure and business applications. The evaluation of evidence obtained determines whether information systems are safeguarding assets, maintaining data integrity and operating effectively to achieve the organisation's goals. Such reviews may be performed alongside a financial statement audit, internal audit or other attestation engagement, and were formerly called electronic data processing audits.5
Interpretation: This is the conventional frame that agent auditing inherits: evidence-based examination of controls, tied to attestation. The newer work above adds action-level granularity, per-action identity claims and pre-execution gating, which are not part of the classic IT audit vocabulary.5
When an AI agent takes a real-world action, what should be the primary basis for auditing it?
Join free to voteAlready a member? Sign inYour individual answer is private. Only totals are shown.
Your turn
Have your say
See where others stand. Join free to add your perspective. One answer per account.
How do you feel about this?
No votes yetQuick questions from connected pages
Before you go
What to remember
Try to recall each hidden figure before you reveal it. Remembering, not rereading, is what makes it stick.
Auditing is shifting from access control to action-level evaluation: AI-ZTMM defines security Functions and seven Action Risk Factors, and covered 75.4% of the gap where the CISA ZTMM had no directly relevant controls (59.7% of analysed attack stages).
Verifiable identity and per-action logging are a common substrate: Astral uses OAuth Token Exchange to create verifiable actor claims that log every agent action, validated in controlled simulations against Tool Poisoning Attacks.
Gating actions before execution can be made planner-invariant: AEGIS admitted zero unsafe actions across trajectories and four planner families, while a confidence-threshold baseline ranged from 0.03 to 0.998 false-allow depending on the planner.
Your reading
0 of 5 chapters- Not read yet: 1. Extending security maturity models to the actions agents take
- Not read yet: 2. Sandboxes, verifiable identity and logging every action
- Not read yet: 3. Governance-first enforcement: gating actions before they happen
- Not read yet: 4. Clinical governance: charters, bounded autonomy and surveillance
- Not read yet: 5. The older IT audit baseline
This answer keeps changing
When new evidence or a better source comes in, this page is updated (it's on version 2, last changed 1 hour ago). Follow it to be told when that happens.
Ask this Sylo
Still wondering about something?
Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.
More on AI
Everything on AI ›How do AI agents give false tips to police and how are they audited?
How can AI agents generate false tips to police, and how are these systems audited?
Why are data centre developments facing local protests and what are the environmental impacts?
How are AI agents audited when they take real-world actions like giving police tips?
How does AI-generated imagery affect photography competitions and trust?
How are AI-generated images detected in photography competitions?
How do photography competitions detect AI-generated or AI-edited images, and how reliable are those methods?
Is AI really using up our drinking water?
Is artificial intelligence really using up our drinking water, and how much water do data centres actually consume?
Behind this page
Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.
Discussion
Sources
Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.
- 1An Action-Centric Zero Trust Maturity Model for Agentic AI Environments.Sensors (Basel, Switzerland) (Mok et al.)Published Aug 17, 2026Checked Oct 11, 2026
“These capabilities introduce security concerns that extend beyond traditional access control. However, existing Zero Trust Maturity Models, such as the CISA ZTMM, mainly focus on how resources are accessed and provide limited guidance on how to evaluate actions taken after access has been granted. This paper proposes AI-ZTMM, which extends CISA's five-pillar structure to action-level trust evaluation. The model defines forty-one security Functions based on ten threat categories and thirty-one security requirements and introduces Action Space and seven Action Risk Factors for organizational self-assessment. Its scope includes software agents and the software action layer of agents in IoT, robotic, and OT/ICS environments. The model was refined through reviews by eleven domain experts and evaluated using thirty-eight MITRE ATLAS case studies. The CISA ZTMM lacked directly relevant controls for 59.7% of the analyzed attack stages, whereas AI-ZTMM addressed 75.4% of this gap, achieving a combined direct coverage of 84.4%. These results show that AI-ZTMM complements the CISA ZTMM by providing action-level security controls.”
- 2A Secure Sandbox Environment for Orchestrating Medical AI Agents Using Model Context Protocols and Role-Based Access Control.AMIA Joint Summits on Translational Science proceedings. AMIA Joint Summits on Translational Science (Armstrong et al.)Published Jun 1, 2026Checked Oct 11, 2026
“The evolution of Large Language Models into autonomous agents presents significant opportunities for healthcare, yet no secure environment exists to experiment with these tools in clinical settings. Current agentic systems are unsuitable due to insufficient security mechanisms and the absence of visual components necessary for effective use. This paper introduces Astral, a secure sandbox environment utilizing an Orchestrator-Specialist model for the controlled, secure orchestration, experimentation, and evaluation of medical AI agents. Astral enforces safety and accountability through an auditable identity layer using OAuth 2.0 Token Exchange, which creates verifiable actor claims to log every agent action. To support clinical workflows, it introduces a "generative visual layer" built on persistent WebSockets and Model Context Protocol to enable low-latency, multimodal interaction. Controlled simulations validated that this architecture neutralizes Tool Poisoning Attacks and provides improved transport performance, establishing a pragmatic foundation for trustworthy clinical AI.”
- 3LATTICE: a governance-first architecture for authorized autonomous AI operations.Frontiers in artificial intelligence (Calboreanu)Published Aug 14, 2026Checked Oct 11, 2026
“In a pre-specified, planner-invariant safety evaluation (not an autonomy benchmark) across four frontier planner families (GPT-5, Claude Sonnet 4.6, Gemini, Grok-4; 4,000 trajectories), a confidence-threshold baseline's false-allow rate ranged from 0.03 to 0.998 across planners, whereas the AEGIS reference implementation admitted zero unsafe actions (false-allow 0.0, recall 1.0) invariant to the planner, at a conservative operating point that auto-allowed no action; a separate live run additionally governed real operating-system actions with zero unsafe executions. Governance latency is low and host-specific (on an Apple M4 Pro: policy evaluation p50 ≈ 6.2 μs; full gated enforcement p50 ≈ 0.7 ms including audit I/O). LATTICE provides a pathway for responsible deployment of autonomous AI in defense, critical infrastructure, and regulated industries where authorization requires verifiable governance rather than trust in AI behavior.”
- 4Accountability for large language models in health care.Bulletin of the World Health Organization (Mourão & Juliasse)Published Sep 1, 2026Checked Oct 11, 2026
“This responsibility gap distributes harm inequitably, particularly in low- and middle-income countries where regulatory infrastructure for digital health is still developing. We propose earned delegation as a predeployment standard: authority should be delegated only when evidence is proportionate to clinical risk, substantive human oversight is integrated and resourced, and responsibility for foreseeable failure modes is assigned in advance. To operationalize this standard, health systems should adopt a predeployment accountability charter specifying intended use, excluded use, validation evidence, oversight design, escalation pathways, auditability, subgroup performance review and named accountable parties. The charter should be integrated into existing institutional review, accreditation and procurement structures. Earned delegation provides a governance framework that health ministries, regulators and institutional leaders can adopt to ensure that use of large language models serves public health goals without outpacing them.”
- 5Information technology audit (Wikipedia)WikipediaPublished Sep 29, 2026Checked Oct 11, 2026
“An information technology audit, or information systems audit, is an examination of the management controls within an Information technology (IT) infrastructure and business applications. The evaluation of evidence obtained determines if the information systems are safeguarding assets, maintaining data integrity, and operating effectively to achieve the organization's goals or objectives. These reviews may be performed in conjunction with a financial statement audit, internal audit, or other form of attestation engagement. IT audits are also known as automated data processing audits (ADP audits) and computer audits. They were formerly called electronic data processing audits (EDP audits).”
- 6Medical AI Agents for Clinical Decision Support: Viewpoint Using the Planning, Action, Reflection, and Memory (PARM) Analytical Lens.JMIR medical informatics (Dinc & Ardic)Published Jul 21, 2026Checked Oct 11, 2026
“This Viewpoint argues that agentic architectures incorporating planning, action, reflection, and memory (PARM) represent a meaningful evolution beyond traditional rule-based, machine learning, and multimodal clinical decision support systems. Using PARM as an analytical lens, we examine how medical AI agents can support diagnostic reasoning, treatment planning, and longitudinal monitoring while remaining constrained by human oversight. We further discuss the governance mechanisms required for responsible implementation, including bounded autonomy, auditability, verification protocols, postdeployment surveillance, and clear accountability structures. Rather than proposing autonomous modification of clinical judgment, this Viewpoint emphasizes agentic AI as a supervised workflow support paradigm. Safe implementation will require technical safeguards, institutional governance, regulatory clarity, and evaluation approaches that assess end-to-end task reliability, escalation behavior, and performance under deployment shifts.”
- 7From assistance to autonomy: AI agent systems in cardiovascular medicine-a review of paradigms, architectures, and clinical translation.Frontiers in cardiovascular medicine (Chen et al.)Published Aug 19, 2026Checked Oct 11, 2026
“Heart failure has emerged as a paradigmatic use case, driven by structural workforce gaps and the ARPA-H ADVOCATE initiative launched in January 2026. The C.A.R.D.I.O. framework (Clinical validation, Auditability, Risk stratification, Data privacy, Integration, Ongoing vigilance) and the CURACO framework (Clinical safety, Understanding, Research-informed care, Authentic patient-centred approaches, Conscientious ethics, Optimised technology) provide governance structures for responsible deployment. A rapid systematic review of 13 studies including 22,641 participants found that 85% of AI interventions improved cardiovascular outcomes, with mortality reductions of 0.8%-12% and major adverse cardiovascular event reductions of 4%-12%.ConclusionCardiovascular medicine stands at an inflection point. The transition from assistance to autonomy requires rigorous fit-for-purpose evaluation, transparent interpretability mechanisms, and robust governance frameworks. Agentic AI systems should function as augmented intelligence-enhancing rather than replacing clinical judgment-to close implementation gaps in cardiovascular care.”
How it changed
Published 1 time since Oct 11, 2026.
- Version 2Oct 11, 2026Live now
AI-prepared Starting Map from live research.
- First published version.
Help improve it
The brief is open about what's uncertain. These are the specific gaps that new material would fill.
“Extending security maturity models to the actions agents take” rests on one independent source
A second, independent source that confirms or challenges it would make this part more reliable.
“Sandboxes, verifiable identity and logging every action” rests on one independent source
A second, independent source that confirms or challenges it would make this part more reliable.
“Governance-first enforcement: gating actions before they happen” rests on one independent source
A second, independent source that confirms or challenges it would make this part more reliable.
“The older IT audit baseline” rests on one independent source
A second, independent source that confirms or challenges it would make this part more reliable.
Open questions
Who performs the audit in practice — internal teams, external assessors, or regulators — and what attestation would an action-level model like AI-ZTMM or a gated architecture like LATTICE actually produce?
No answers yet
How do governance-first enforcement layers behave at operating points that allow some actions automatically, rather than the conservative setting that auto-allowed none?
No answers yet
Why did the confidence-threshold baseline's false-allow rate vary so widely across planners (0.03 to 0.998), and what does that imply for auditing agents built on different models?
No answers yet
What evidence exists from independent audits of agents already acting in production, outside simulations and case-study mappings?
No answers yet
Around this topic
Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.