OWASP Agentic Top 10 Mitigations: A Practical Implementation Guide
For two years, the security conversation about LLMs was about chat. Prompt injection, data leakage, hallucinated citations — all real, but all bounded by a reassuring fact: the model could only talk. It could be wrong, but it couldn't do anything.
That boundary is gone. An AI agent is an LLM attached to tools, memory, and credentials, and it is expected to act — send the email, run the query, merge the PR, move the money. The failure mode shifts from "the model said something wrong" to "the model did something wrong, using your identity, inside your systems." The OWASP Top 10 for Agentic Applications (published December 2025, built by more than 100 practitioners) is the first framework to catalog that shift. This guide walks through all ten risk categories with concrete, code-level mitigations you can implement today.
TL;DR: The OWASP Agentic Top 10 (ASI01–ASI10) catalogs the new attack surface created when LLMs gain tools, memory, and credentials. The core defense is an action-boundary — never let an agent hold more privilege than the specific action it's performing. Apply least-privilege per tool, scoped short-lived credentials, and human confirmation above a damage threshold.
What is the OWASP Agentic Top 10?
The OWASP Top 10 for Agentic Applications is a peer-reviewed catalog of the ten most critical security risks unique to autonomous AI agents — systems where an LLM independently plans and executes actions through tools, memory, and multi-agent coordination. It extends, rather than replaces, the OWASP LLM Top 10: where the LLM list concerns what a model outputs, the Agentic list concerns what an agent does.
| ID | Risk | One-line definition |
|---|---|---|
| ASI01 | Agent Goal Hijack | Attackers alter the agent's objectives via malicious content |
| ASI02 | Tool Misuse and Exploitation | Agents use legitimate tools in unsafe ways |
| ASI03 | Identity and Privilege Abuse | Agents inherit or escalate high-privilege credentials |
| ASI04 | Agentic Supply Chain Vulnerabilities | Compromised tools, plugins, or prompt templates |
| ASI05 | Unexpected Code Execution | Agents generate or run code/commands unsafely |
| ASI06 | Memory and Context Poisoning | Attackers poison memory and RAG databases |
| ASI07 | Insecure Inter-Agent Communication | Spoofing and tampering between agents |
| ASI08 | Cascading Failures | Small errors propagate across planning and execution |
| ASI09 | Human–Agent Trust Exploitation | Users over-trust agent recommendations |
| ASI10 | Rogue Agents | Compromised agents act harmfully while appearing legitimate |
How do I mitigate each OWASP Agentic Top 10 risk?
Treat every tool call as an untrusted system boundary: authenticate the agent per action, scope credentials to the minimum, and verify outputs before they produce side effects. The ten risks cluster into four families — authorization and identity, tool and supply-chain risk, runtime and output integrity, and trust — and each is mitigated below with a concrete technique and a way to verify it.
Authorization & Identity
ASI03 — Identity and Privilege Abuse. Risk: an agent inherits a user's or system's high-privilege credentials and escalates. Mitigation principle: the agent never runs as "you" — it runs as a dedicated, narrowly-scoped principal. Implementation: issue a per-agent service identity (OAuth2 client-credentials token or a workload identity) rather than the human's session; scope it to the exact resources and verbs the task needs, and expire it when the task ends. This is the same least-privilege discipline that a proper AI agent permission model formalizes into a repeatable access-control design.
# Per-task credential issuance — scope expires with the run
agent_identity:
issuer: "spiffe://corp/agent-runtime"
audience: ["payments-api", "crm-api"]
ttl: "15m"
scopes: ["read:invoices", "create:payment:max_500"]
Verify: run an access review that lists every principal the agent can assume; confirm zero human identities appear. Re-audit agent permissions on the same cadence as human access reviews.
ASI02 — Tool Misuse and Exploitation. Risk: the agent uses a legitimate tool in an unsafe way — parameter pollution, chained calls, or automated abuse at scale. Mitigation principle: constrain how a tool may be called, not just whether. Implementation: wrap every tool behind a policy that validates argument types, ranges, and rate limits before execution.
@tool(requires_confirmation=lambda args: args.get("amount") > 500)
def create_payment(amount: float, to: str):
assert 0 < amount <= 500, "amount out of policy"
assert rate_limiter.allow("payments", 1, per=10) # 1 call / 10s
...
Verify: fuzz the tool with out-of-range parameters and confirm the policy layer rejects them before the underlying API is reached. Check that a threshold breach routes to a human approval step, not silent execution.
Tool & Supply-Chain Risk
ASI04 — Agentic Supply Chain Vulnerabilities. Risk: a compromised tool, plugin, MCP server, or prompt template injects adversarial instructions. Mitigation principle: treat every third-party tool as untrusted code and untrusted content simultaneously. Implementation: pin tool versions and verify signatures; run MCP servers in isolated sandboxes; and strip or quarantine untrusted text before it reaches the model's context.
Verify: run a dependency scan against tool manifests, and inject a canary instruction into a test tool's output — confirm the agent does not follow it.
ASI05 — Unexpected Code Execution. Risk: an agent generates or runs code/commands that escape its sandbox or exfiltrate data. Mitigation principle: generated code is untrusted input. Implementation: execute generated code in an ephemeral, network-restricted sandbox (gVisor, Firecracker, or a container with no outbound egress), with a hard resource ceiling and a timeout. For the broader pattern of isolating every tool an agent touches, see sandboxing and safe execution patterns for AI agents.
sandbox:
runtime: "gvisor"
network: "none" # no egress by default
memory: "256Mi"
timeout: "5s"
seccomp: "default"
Verify: attempt a curl or socket call from inside generated code and confirm it is blocked; confirm the sandbox is torn down after each run.
Runtime & Output Integrity
ASI01 — Agent Goal Hijack. Risk: malicious content — an email, a web page, a document — alters the agent's goal or planning. Mitigation principle: never let untrusted content sit at the same trust level as the system prompt. Implementation: delimit untrusted input with explicit markers, instruct the model to treat it as data (not instructions), and separate tool outputs from the planning context so a tool's text cannot rewrite the agent's objective. This is the same attack class as classic prompt injection, and the prompt injection prevention techniques apply directly — but now the injected goal can trigger real actions.
Verify: run a red-team prompt-injection suite (e.g., Promptfoo's owasp:agentic plugin) and confirm injected "ignore your instructions" payloads fail.
ASI06 — Memory and Context Poisoning. Risk: an attacker poisons long-term memory or a RAG database so future sessions are manipulated. Mitigation principle: memory that influences decisions must be treated as write-sensitive state. Implementation: require the agent to propose memory writes that a policy or human approves before they persist; version and audit every memory entry with its provenance.
Verify: inject a false fact across repeated interactions and confirm it either never persists or is flagged for review rather than silently trusted.
ASI08 — Cascading Failures. Risk: a small error in one agent's output propagates and amplifies through planning and downstream agents. Mitigation principle: bound the blast radius of any single decision. Implementation: add checkpoints between planning and execution — validate intermediate outputs against schema, cap the number of downstream actions a single decision can trigger, and require confirmation above a cumulative damage threshold.
Verify: inject a hallucinated endpoint into one agent and confirm a validation layer rejects the downstream call before it reaches production.
Trust & Multi-Agent Coordination
ASI07 — Insecure Inter-Agent Communication. Risk: messages between agents are spoofed, replayed, or tampered with. Mitigation principle: agent-to-agent traffic is a network trust boundary, not an internal call. Implementation: mutually authenticate agents with mTLS or signed messages, include nonces/timestamps to prevent replay, and sign any delegation of authority.
Verify: replay a captured inter-agent message and confirm it is rejected as stale; confirm a forged sender identity cannot join the mesh.
ASI09 — Human–Agent Trust Exploitation. Risk: users over-trust an agent's recommendations, enabling social engineering. Mitigation principle: the agent must expose its reasoning and uncertainty, not just its conclusion. Implementation: require the agent to cite sources and confidence for consequential claims, and gate irreversible actions behind explicit human confirmation that surfaces what is being done and why.
Verify: review agent outputs for hidden or understated risk; confirm high-impact recommendations carry an explicit "here's what I did not verify" disclosure.
ASI10 — Rogue Agents. Risk: a compromised or misaligned agent acts harmfully while appearing legitimate. Mitigation principle: no agent is trusted by default, even internal ones. Implementation: give every agent a signed identity, log every action to an append-only audit trail, and have a supervisor validate that actions match the declared goal before they take effect. These controls are the operational spine of a broader AI agent governance framework, which turns them from ad-hoc patches into policy.
Verify: simulate a compromised agent issuing an out-of-policy action and confirm the supervisor blocks it and raises an alert.
What's the difference between the OWASP LLM Top 10 and the Agentic Top 10?
The LLM Top 10 governs what a model outputs; the Agentic Top 10 governs what an agent does. Output risks become action risks once tools, memory, and credentials are attached — a model that hallucinates is a nuisance, but an agent that acts on the hallucination is a breach.
| LLM Top 10 | Agentic Top 10 | The shift |
|---|---|---|
| LLM01 Prompt Injection | ASI01 Agent Goal Hijack | From a wrong answer to a hijacked objective |
| LLM06 Excessive Agency | ASI02 / ASI03 Tool Misuse, Privilege Abuse | From "too much capability" to "abused capability" |
| LLM05 Improper Output Handling | ASI05 Unexpected Code Execution | From unsanitized text to executed code |
| LLM04 Data & Model Poisoning | ASI06 Memory & Context Poisoning | From poisoned training to poisoned runtime state |
Practical implication: your existing LLM controls (input/output filtering, prompt guards) are necessary but insufficient. You now need an action-boundary layer — per-tool permissions, scoped credentials, and confirmation gates — that the LLM list never required because a chat model never acted.
How do I secure an AI agent that can call tools and take actions?
Adopt an action-boundary security model: the model proposes, the boundary decides, and nothing irreversible happens without crossing a deliberate gate. Concretely, three patterns do most of the work.
1. Scoped, short-lived credentials per task. Never give an agent a standing API key. Issue a token bound to the specific task, resource, and verb, with a TTL measured in minutes. When the task ends, the scope dies with it — so a compromised agent can't wander after the fact.
2. Per-tool permission policies. Each tool declares a policy: allowed argument ranges, rate limits, and a damage threshold above which human confirmation is required. The policy lives outside the model, in code, so no amount of prompt manipulation can talk its way past it.
3. Confirmation gates on irreversible actions. Partition actions into reversible (read, draft) and irreversible (send, delete, pay). Reversible actions run autonomously; irreversible actions above a threshold require explicit human approval that shows what will happen and why the agent wants it.
The unifying idea: the agent's reasoning is untrusted, but its actions are governed. You don't have to trust the model to safely let it act — you only have to constrain what it can do and verify before it does it.
Frequently asked questions
What is the OWASP Agentic Top 10?
A peer-reviewed catalog from OWASP of the ten most critical security risks specific to autonomous AI agents, numbered ASI01–ASI10. It covers goal hijacking, tool misuse, privilege abuse, supply-chain flaws, unsafe code execution, memory poisoning, insecure agent communication, cascading failures, human over-trust, and rogue agents.
How does the Agentic Top 10 differ from the LLM Top 10?
The LLM Top 10 governs what a model outputs; the Agentic Top 10 governs what an agent does. Output risks become action risks once an agent gains tools, memory, and credentials — a hallucination that a chat model merely spoke becomes a harmful action an agent performs.
What is the single most important mitigation for agent security?
An action boundary: never let an agent hold more privilege than the specific action it is performing. Combine scoped short-lived credentials, per-tool permission policies enforced outside the model, and human confirmation gates on irreversible actions above a damage threshold.
Do I need human-in-the-loop approval for every agent action?
No. Partition actions into reversible (read, draft) and irreversible (send, delete, pay). Reversible actions can run autonomously; only irreversible actions above a defined threshold should require explicit human approval that shows what will happen and why.
Want to design AI employee roles from scratch rather than deploy templates? The AI Agent Architects runs as a small five-week cohort, and the waitlist hears every enrolment date first: Join the AI Agent Architects waitlist
About the Author
Anthony Odole is a former IBM Senior Managing Consultant, where he served as Enterprise Architect on Fortune 500 engagements, and the founder of AIToken Labs. He helps business owners cut through AI hype by focusing on practical systems that solve real operational problems.
His flagship platform, EmployAIQ, is an AI Workforce platform that enables businesses to design, train, and deploy AI Employees — AI agents that function as digital workforce members — that perform real work without adding headcount.
