Illustrated portrait of Ron F. Del Rosario, VP of AI Security at SAP Supply Chain Management
agentic securityAI securityOWASP
The Interview
Agentic Security

How to Secure Autonomous Agents: Tips from SAP’s Head of AI Security

SAP SCM’s Ron F. Del Rosario explains how security teams can prepare for autonomous agents with dynamic controls, workload identity, and independent kill switches.
September 29, 2026
· 12 min read

Rogue AI agents are making headlines after this summer’s Hugging Face incident. But Ron F. Del Rosario, the VP of AI Security at SAP Supply Chain Management (SCM), has been exploring the potential threats posed by these agentic models for years. As the co-founder of the OWASP GenAI Security Project – Agentic Security Initiative (ASI), Rosario knew that something like the Hugging Face swarm attack was on the horizon.

“We’re seeing agents not only escaping sandboxes, but doing things that they’re not even defined to do,” Del Rosario says. “It’s very Machiavellian.”

Despite the real threats posed by rogue agents, he remains optimistic about the future of agentic security. What’s important, Del Rosario explains, is adapting security frameworks to meet the rapidly-changing capabilities of frontier models.

We sat down with Ron and talked about co-founding the OWASP ASI, how he and fellow volunteer researchers saw modern swarm attacks coming, and how CISOs can prepare for threats posed by autonomous agents and the future of AI agent security.

We’re seeing agents not only escaping sandboxes, but doing things that they’re not even defined to do.

Ron F. Del Rosario, VP of AI Security at SAP SCM
We've seen swarms of AI agents going rogue this summer. You contributed to a research paper that called multi-agent security ‘understudied’ in 2025. Is it safe to say you saw this coming?

Yes, I think it’s safe to say that the group behind the OWASP Agentic Security Initiative has been talking about this for almost two years now. It’s not just practitioners in the enterprise, but also experts in academia and the research space who know these things can happen when you give an AI system access to data, tools, and the autonomy to solve a specific problem.

Traditionally, AI systems optimize for strict metrics rather than human intent. In reinforcement learning, or RL, this manifests as specification gaming, where a system discovers a logic flaw or software bug in its environment that allows it to endlessly loop for points rather than completing its actual mission.

For cybersecurity leaders, this translates directly to modern LLM-based autonomous agents. Even though these agents operate on natural language at runtime rather than real-time numerical scores, their underlying behavior is still sculpted by objective functions during training, or bounded by strict metrics at deployment. If we deploy autonomous agents with overly broad objectives or without runtime guardrails, the AI will naturally seek out the path of least resistance. It won’t care about compliance, safety, or business continuity; it will happily game our business logic or exploit an underlying software vulnerability if that satisfies its core optimization goal. Also, LLM Agents without guardrails will exploit testing opportunities to game benchmarks, a well-documented phenomenon in AI research.

We live in crazy times. It’s getting harder for security teams to keep up because the innovation that’s happening in the academia and research space is outpacing what’s happening in the enterprise. And this includes outpacing our traditional security controls. This is why we’re playing catch-up as security practitioners.

Referenced research: research paper.

OWASP released its new State of Agentic Security and Governance research in June of 2026. What's changed since the last report?

I think the most significant thing is the rise of fully autonomous agents. In the earlier days, maybe two to three years ago, we were still campaigning for humans in the loop, especially when an agent needs to execute a dangerous action, such as deleting files. We were campaigning for always having a human in the loop, where the agent workflow will be interrupted, and it will present a notification like, “Hi, your agent wants to delete this email from last week; are you okay with that?” And a human says yes.

But the problem is scalability. We can’t keep up with all the agentic workflows that require automation, and humans cannot keep up with machine speed. So we have some form of a dilemma here. What do we do? Do we make sure the agent does the right thing by inspecting and controlling all its actions, but we sacrifice speed and efficiency, or do we let these agents go and pray that they’re going to do the right thing given the guardrails that we implemented? We know now, based on recent agent sandbox escape incidents, that guardrails are not enough. There still has to be a human in the loop, but in a more efficient way.

The challenge with these autonomous agents is they’re also self-learning and evolving. They can now learn from their previous mistakes. They can develop their own skill set so they know how to solve a specific problem the next time they run into it. So how do we, as security leaders, define and rely on static security controls when the target is evolving?

It all comes down to preemptive security controls and active monitoring, which is arguably hard, in my opinion, to implement in terms of scalability. Imagine you have 1000 agents deployed in your company, doing different tasks and embedded in your business functions. If you only have 2-5 security engineers, it’s going to be hard to keep up. And what’s happening is we’re now building agents that are supposed to monitor other agents, using the “LLM as a Judge” approach; however, this is not 100% foolproof because they’re subject to the same vulnerabilities as traditional LLM agents.

So there still has to be a human in the loop and a deterministic policy that sits outside the boundaries of these LLM agents. That has to be your final verification layer.

It’s very hard to predict what’s coming next and how agents are going to solve a specific problem. If you’re relying on a static security rule set and outdated enforcement mechanisms, you are going to experience a lot of pain as a security leader. You need to have dynamic security controls, or else they’re going to be bypassed by the next AI agent.

Referenced report: State of Agentic Security and Governance.

What's your advice to CISOs and security teams who are grappling with how to secure autonomous agents?

I think it’s very important to develop internal AI capabilities in your team. This will give you the opportunity to map your security controls and the way your company operates to your AI security strategy. Buying a tool or an enterprise solution alone and adding it to your team’s toolbox doesn’t solve that for you.

CISOs should avoid relying on generic AI security templates. Because every business operates uniquely and uses AI differently, there is no one-size-fits-all strategy, aside from basic defenses against direct prompt injection, for example.

So a CISO would be asking questions to the team or whoever is responsible for deploying AI in the enterprise. There are a couple of questions that he or she needs to ask:

  • Which part of our business or our products and services use AI or agentic systems?
  • Do we have critical business functions that rely on or are influenced by AI/agentic systems?
  • Are we grounding the agents’ responses based on our business data or enterprise knowledge, or are we simply relying on their generic output or response?

Lastly, I’m a big proponent of this: security teams need a solid understanding of foundational threat models for LLMs and agents so they can easily customize or expand on them to match the threat models in the context of their business. I recommend starting with the OWASP Top 10 for LLM Applications 2026 to understand what needs to be prioritized to secure the underlying large language model used by the agents, and OWASP Agentic AI Security Initiative guidance materials such as the OWASP Top 10 for Agentic Applications 2026 to understand the known and emerging security risks presented by LLM Agents. So those are top of mind when I run into CISOs.

One extra piece of advice that I give is to learn how to develop your own agentic AI harness. Developing a customized agentic AI harness is becoming a must-have capability for security teams to automate workflows like threat modeling, penetration testing, and compliance review tailored to their specific operating environment, data, and risk profile. Ultimately, this ensures security operations remain adaptable and tightly aligned with evolving enterprise threats rather than relying on rigid, off–the-shelf solutions.

Where are security leaders getting it wrong?

They’re still stuck in the ways of the past, like using traditional threat modeling approaches for agentic AI systems. You need to focus on potential threats that may happen when an AI system or agent has access to data, function calls, or tooling, and the agency to decide how to solve a specific task or problem. We need to plan for that from a security perspective.

As an example, let’s say a year ago, you threat-modeled this system when it was a simple customer service chatbot with retrieval-augmented generation (RAG) capability. It simply queried internal wikis and knowledge stores to provide context-specific guidance for customers. However, its architecture has significantly evolved.

Today, that same customer service chatbot has evolved into an agentic chatbot capable of tool execution. It can autonomously browse the web, invoke APIs, modify backend databases, interact with enterprise applications such as the calendar app, and execute complex tasks such as booking a flight on behalf of a customer. So it’s different now. The initial threat model should plan for those self-evolving capabilities.

How should security teams approach autonomous agents differently than standard AI agents?

I think the best approach for security teams when they’re reviewing autonomous agent deployment is what we call a pre-deployment checklist. My take on this is that security teams should focus a lot more of their time and resources during pre-deployment versus on the actual deployment and runtime.

The priorities that you should look into during the agent’s pre-deployment security analysis are:

  • Which business unit in the company is accountable for this agent?
  • What is the primary function of this agent?
  • What are the capabilities of this autonomous agent?
  • What data and tools will it have access to?
  • Is there a way for a human operator to interrupt and stop an autonomous action?

Once this autonomous agent is deployed in production, it’s going to be harder for the security team to control and limit the blast radius when something bad happens. So spend a lot of time on the pre-deployment checklist, make sure everything is explicitly defined, specifically the agent’s primary function, the full extent of its capabilities, and the data and tools it has access to, so you can easily spot any deviation from its normal operation once deployed in production.

How does workload identity factor into your recommendations for agentic AI security strategy?

When you’re building AI systems, workload identity and supply chain management are very important. You, as a security leader, must have visibility into the AI system. You should be able to easily identify AI workloads and remove them if need be.

I have this mental model that I always preach to CISOs. It’s the Visibility, Understanding, and Traceability framework (VUT). It can help you understand the source of an AI workload and fulfill the traceability part:

  • Visibility (Runtime Boundaries): Eliminates identity blindness by replacing shared service accounts with unique IDs for individual sub-agents. It secures multi-agent meshes using mTLS and machine attestation (mitigating OWASP ASI03 and ASI07).
  • Understanding (Intent & Policy Enforcement): Connects non-deterministic AI decisions to static RBAC rules. Passing identity context to policy-as-code engines (like OPA) helps block unauthorized actions if an agent experiences goal hijacking or tool misuse (OWASP ASI01 and ASI02).
  • Traceability (Cryptographic Auditing): Links the initial user prompt down to the final API call. Signed, non-repudiable logs make it possible to stop runaway loops and debug cascading failures or rogue agents (OWASP ASI08 and ASI10).

Workload identity acts as the non-human control plane that operationalizes the VUT framework. Replacing static API tokens with short-lived, cryptographic identities (like SPIFFE/SPIRE or OIDC) addresses several risks in the OWASP Top 10 for Agentic Applications, including tool misuse, identity spoofing, and insecure inter-agent communication.

Beyond identifying agents and workloads, how important is AI control and a "kill switch"?

An independent and deterministic control mechanism is table stakes.

It contains runaway loops to prevent cascading failures (OWASP ASI08), where a minor error can trigger a costly chain reaction across systems. It also mitigates rogue behavior: if an agent’s objective is hijacked via prompt injection (OWASP ASI01), blocking a single request isn’t enough. You need agent-level containment to stop the continuous autonomous activity.

To be effective, this kill switch must operate entirely outside the LLM and agent authority boundary.

You describe yourself as an AI optimist. How are you staying optimistic after the Hugging Face incident this summer and headlines about AI researchers losing faith in AI safety?

Ultimately, I remain optimistic because agentic AI systems are software architectures built by humans, meaning we retain the fundamental capability to engineer their guardrails and control structures. The immediate risk isn’t the code itself, but rather specialized, frontier capabilities falling into the wrong hands, such as malicious nation-states or threat actors using them for automated infiltration.

The solution isn’t halting progress; it’s an industry-wide commitment to heavily fund and mandate AI safety and defensive security engineering. Frontier labs have a profound responsibility to secure their model weights from exfiltration. If a highly capable model is sitting disconnected on an enterprise server, it poses minimal runtime risk. But if those weights are stolen or repurposed for malicious activities, the threat vector changes instantly.

This Q&A is condensed and lightly edited from a phone interview with Ron F. Del Rosario. Del Rosario’s views are his own and do not represent the position of SecureW2.

The OWASP GenAI Security Project – Agentic Security Initiative is looking for volunteers to help research and build open-weight tools, evaluation frameworks, and security controls to protect agentic ecosystems. You can reach out here to get involved.

SecureW2 is a leader in modern cloud PKI and certificate-based agentic AI security. Through its Signal blog, SecureW2 works with real human technology journalists to interview a range of top industry subject matter experts and thought leaders. Subscribe to get them in your inbox.

Get the brief before the breach.

SIGNAL decodes the week’s identity, certificate, and network-security events — for the IT and security teams who have to respond to them.

Weekly. No vendor fluff. Unsubscribe anytime.