The Exhausting Paradox of AI Security

By Lee Rossey, CTO and Co-Founder SimSpace [ Join Cybersecurity Insiders ]

In July, the world was startled by the news that OpenAI’s security testing agents had escaped their sandbox and hacked Hugging Face. During a security test, an agent hit a dead end, found its own workaround to get online anyway, and stumbled into a shared file system where it found it had write access. It then shared that information with other agents, which began coordinating more aggressively to pursue external infrastructure they believed might hold answers to their tasks, culminating in the compromise of Hugging Face. Just as everyone was losing their heads over what it all meant, Anthropic, not to be outdone, suddenly announced that their agents had also managed to escape from a sandboxed environment and hacked three companies. 

While industry pundits launched into various hot takes, I couldn’t help thinking about how those making AI procurement decisions really felt about this latest development; it hardly engenders trust that these models can be controlled. At the same time, security leaders were having to answer tough questions about what defenses they have in place to meet this unexpected threat. This is the exhausting paradox of AI – everyone wants to have the most powerful models in their arsenal, but they also want to know they can trust it not to cause harm to their company or customers. 

Beyond the headlines and hype, enterprises are looking for reliable AI solutions that can help them solve their problems. Meanwhile, AI agents are becoming increasingly autonomous, and the more autonomy, the greater degree of testing needed to prove they’re good. With these frontier model attacks adding a new element of risk, organizations need a safe place to evaluate whether their controls, detection, containment, and teams can withstand it.

Right now, security organizations are facing three main problems when it comes to deploying AI confidently. For those buying AI solutions, they need to know how it behaves in their environment, against their traffic, under real load. For those building their own AI models, they need to know whether their training data is representative enough for the model to generalize beyond the attacks it was trained on. Finally, if organizations are trying to run AI and human analysts together, they need to know where their coordination breaks down before an incident forces them to find out. 

And as the OpenAI/Hugging Face example showed us, AI models are not reliably predictable. Before AI, security software was largely deterministic. An attacker could subvert the logic, but the system still did what it was told. With large language models, however, the same input can produce different outputs on different runs. Deciding whether an alert is real or noise is already hard for experienced analysts, so asking for an AI to do it reliably adds another layer of risk. Beyond the autonomous rule-breaking we’ve already seen, other failure modes to watch for include hallucinations, prompt injection vulnerabilities, context blindness, and model drift. If OpenAI and Anthropic are still working through their reliability problems, what hope does anyone else have?

This is why we started the AI Proving Ground Consortium (AIPGC), a group dedicated to helping enterprises train, validate, and prove their AI systems in a controlled environment before those systems ever touch a live security operation. Together, the consortium gives customers a full closed loop for testing AI systems: a realistic environment to run them in, network-level evidence of what they actually do, a live AI agent to test against, an adversary simulating real attacks, security testing of the AI itself, and independent analysis to check the results. The result is that no single vendor is grading its own work. A customer gets to see how an AI system holds up end to end, from the range it’s tested in to the attacks thrown at it, to the verdict on whether it can be trusted, instead of piecing that together from separate tools.

The consortium works because no single vendor can pull this off alone. A real enterprise runs on dozens, if not hundreds, of vendors. AI security cuts across data, telemetry, agents, workflows, validation, and governance, so assessing one piece and calling it done misses most of the picture. When an AI agent fails, it rarely fails at one point. That’s why trust must be tested across the entire system, not tool by tool. Realistic protection also starts with realistic data. Training and validation only mean something if the environment behind them mirrors the mess of a real network. Skip that step, and AI readiness turns into a box you check once instead of something you keep testing.

It’s that continuous element that’s crucial. AI evolves daily; neither the OpenAI nor Anthropic agent attacks were planned, and that’s exactly why testing for something like a Mythos zero-day matters. It forces companies to find out, ahead of time, whether their controls, detection, containment, and teams can hold up once an agent starts acting on its own. Companies putting AI agents into security workflows need to go beyond verifying the technology works; they need to know it can be trusted under real deployment pressure, tested repeatedly as conditions shift. Threats, environments, and models are never going to stop moving, so testing can’t be a one-time check either.

These attacks point to where this is heading. It won’t be a single AI agent running the SOC. There will be several agents working on investigation, response, threat hunting, and decision support alongside human analysts. That’s why we built AIPGC, and it’s why we’re encouraging security leaders across major industries to start asking how they’d prove AI readiness before it ever touches production. The organizations that win with AI cybersecurity will not be first adopters, but those that understand how to trust it ​​responsibly.

 

Join our LinkedIn group Information Security Community!

No posts to display