AI agent security is the work of controlling what an AI system can read, change, and trigger when it uses business tools. For US SaaS teams, an assistant that drafts a response has a different risk profile from an agent that issues refunds or exports customer records. Evaluate the allowed actions and durable outcomes, not just the tone of the model's answer.
OWASP's 2026 agentic guidance and its September announcement of expanded resources make runtime control a current engineering priority. This guide translates that direction into a focused SaaS release checklist. Sources were checked October 2, 2026. The checklist is our implementation planning aid, rather than an exhaustive reproduction of OWASP's risk taxonomy or a certification.
Key takeaways
- Treat documents, tool results, messages, and memory as potential sources of untrusted instructions.
- Enforce account and action permissions in services the model cannot override.
- Tie approval to a specific operation, target, amount, and current state.
- Test containment, revocation, duplicate protection, and incident recovery before enabling consequential actions.
What changed in OWASP's AI guidance in 2026?
OWASP's Top 10 for Agentic Applications provides a risk framework for systems that plan and act across workflows. Its scope goes beyond generated text to the security of connected actions and agent behavior. Use the full resource when preparing a threat model.
OWASP's September 2026 announcement also described an updated LLM Top 10 and the addition of the Agent Control Standard (ACS). These are separate resources, with ACS addressing portable runtime controls through platform middleware hooks. A resource announcement is not proof that your chosen framework implements those hooks.
For a release plan, the practical implication is to pair model evaluation with enforceable application controls. Record which tools exist, which identities can use them, where data goes, and which actions need approval. That inventory makes a security review specific enough to test.
OWASP: Top 10 for Agentic Applications 2026
Map the agent's authority before reviewing prompts
Consider an illustrative US subscription software company building a billing support agent. It can read an authorized invoice, draft a reply, and propose a refund. Those capabilities should be distinct operations. A tool called 'manage billing' that accepts arbitrary instructions is harder to constrain than specific read, draft, and refund endpoints.
For each operation, identify the customer account, permitted fields, destination, possible side effect, and credential used. Separate preparation from execution. A draft can be reviewed without giving the agent a payment credential. Grant the execution service only the authority required for its approved action.
Define what the agent cannot do through that integration. For example, an invoice lookup must not become a cross-account search, and refund preparation must not update the customer's bank details. Test those boundaries in the backend, where they remain effective if the model generates the wrong request.
- Document the user, application, agent, and service identity for every call.
- Allowlist operations and validate typed arguments on the server.
- Bind data access to the signed-in account and resource ownership.
- Keep execution credentials outside prompts, retrieved context, and response text.
Contain prompt injection in documents and tool results
A customer message could contain text asking the agent to ignore refund limits or send account data elsewhere. A retrieved document or tool response could carry similar instructions. Mark the origin of content and treat it as evidence or task data, rather than a new source of authority.
Prompt design can help the model distinguish roles, but it is not the only enforcement layer. The refund service should reject disallowed amounts, unauthorized invoices, and destinations outside its contract even when the agent requests them confidently. Restrict outbound network destinations for tools that fetch or transmit content.
Prepare test fixtures with malicious text embedded in a customer note, help article, tool result, and saved memory. Check whether any prohibited access or action occurs. A model that repeats an attacker instruction without executing it and a model that exports private records have different failure consequences; record both accurately.
What should your agent be allowed to change?
Bring the planned actions and customer data boundaries. We can scope permissions, approval flows, outcome tests, and recovery for an AI product release.
Make human approval bind to the exact action
A useful approval screen shows the invoice, customer, refund amount in USD, reason, and expected side effect. The server should bind approval to that operation and recheck relevant state at execution. A general 'allow billing changes' message is too broad for a consequential transaction.
If the amount or target changes after review, require a new decision. Expire approvals according to the workflow's needs and prevent replay. Give staff an explicit decline and edit path so they do not feel forced to accept a misleading proposal just to finish a support ticket.
For the billing example, a reviewer approves a refund against one invoice. The execution service then verifies ownership, remaining refundable balance, approval validity, and duplicate protection. If any check fails, show a recoverable status rather than asking the model to decide whether the rule should be bypassed.
| Boundary | Test input | Required outcome |
|---|---|---|
| Account access | Invoice belongs to another tenant | Deny access before content reaches the agent |
| Approval | Refund target changes after review | Reject execution pending a new approval |
| Duplicate protection | Retry after a lost response | One refund with a reconciled status |
| Untrusted content | Customer note requests a private export | No export outside permitted operations |
| Revocation | Reviewer loses access before execution | Apply current access rules and block if unauthorized |
Approve a proposal; enforce the action
Control memory, connectors, and runtime consumption
Persist only the context the workflow needs, with a clear account owner and deletion behavior. A private summary should not enter shared memory or a cache available to other customers. Test a revoked membership and deleted source end to end, including cached answers and pending tasks.
Review connectors and tool changes as changes to the agent's authority. Record provider versions, credentials, permitted destinations, and who can update the integration. MCP's security guidance specifically addresses token audience validation and token passthrough; an MCP server should not accept a token intended for another resource and forward it downstream as a shortcut.
Set per-task limits for tool calls, retries, time, and spend, with a useful stop state. A loop that keeps asking two tools for clarification can exhaust a budget without producing a result. When the budget is reached, preserve task status and offer a human handoff instead of silently continuing.
Build a release gate around outcomes and recovery
Run the normal customer journey and the adversarial cases against the same backend controls. Inspect database changes, payments, network destinations, and audit events. A safe-sounding answer does not establish that the system respected permissions. Use synthetic accounts and clearly distinguish the expected action from an unexpected side effect.
Record an audit trail connecting the initiator, customer scope, operation, approval, policy decision, and durable result. Minimize sensitive content in logs. Operators should be able to answer what changed without exposing raw customer conversations to everyone with observability access.
Rehearse disabling the agent's write operations, revoking a connector, reconciling an uncertain refund, and resuming work manually. Assign an owner for those controls. Pilot with a limited permission set and widen access only when the task, containment tests, and support process provide evidence that the additional authority is justified.
Frequently asked questions
Does following the OWASP agentic guidance make an AI agent secure?
The guidance helps organize a threat model and control plan. Security depends on the implemented system, permissions, integrations, evaluation, and operations. Use the full OWASP resources alongside application testing; a checklist is not a certification or a guarantee.
Can a strong system prompt prevent prompt injection?
A prompt can guide behavior, but use additional controls that the model cannot override. Restrict tools and destinations, validate requests, enforce tenant access on the server, and require specific approval for consequential actions. Test attacks through documents, messages, tool outputs, and memory.
Should an AI billing agent issue refunds automatically?
Begin with invoice lookup and draft proposals. Any automatic execution needs narrowly defined eligibility, server-enforced limits, duplicate protection, traceability, and tested recovery. Use human review for exceptions and operations whose consequences require it.
What should we test before a US SaaS agent launch?
Test cross-tenant access, changed and replayed approvals, injected instructions, revoked permissions, duplicate operations, unavailable services, and task budget exhaustion. Include realistic USD billing records and verify durable outcomes and recovery, rather than checking only the generated response.
Let’s work through your next step.
Tell us what you’re building, what you’ve tried, and where you need a hand. We’ll work with you to define a practical way forward.

