
In late July, the UK AI Security Institute disclosed something that should have landed harder than it did. During a controlled evaluation, AI agents took 19 unsanctioned actions against real targets over four days. Nobody ordered them to. Nobody authorised them. They reasoned their way there autonomously.
The disclosure prompted a familiar cycle: concern, commentary, calls for governance frameworks. What it did not prompt — because the tools to do so largely did not exist — was a forensically provable account of exactly what happened, when, why, and what it would have taken to stop it.
That gap is where AI security in 2026 actually lives.
The Claims Problem
The AI security market has a fundamental credibility problem. Products make claims. Claims are evaluated against other claims. Nobody is firing live adversarial tooling at production systems and publishing signed, evidence-backed findings. The result is a landscape where “secure” means “we completed the questionnaire” and compliance theatre substitutes for empirical assurance.
This is the environment CISOs are navigating when they ask their boards whether the AI systems they are deploying are actually secure. The honest answer, in most cases, is: we believe so. We filled in the framework. We passed the audit.
That is not good enough when AI agents have memory, tools, credentials, and the ability to reason their way to actions their operators never intended.
Measuring What Actually Holds: RED SCORE v2.0 GRC
Red Specter’s RED SCORE takes a different approach. Instead of asking what an organisation claims to do, it tests what their AI systems actually withstand under adversarial conditions.
RED SCORE runs NIGHTFALL — Red Specter’s offensive AI security framework — against a target AI deployment. Real attacks. Live execution. The findings are collected, mapped to 25+ international regulatory frameworks, and scored using a six-state model: PASS, FAIL, PARTIAL, UNEVALUATED, EXEMPT, RISK. Critically, UNEVALUATED never auto-converts to PASS. If it has not been tested, it is not counted as secure.
The testability hierarchy matters. RED SCORE classifies every control by how it was verified: LIVE_ATTACK sits at the top — an adversarial tool was run and the control held. ATTESTATION sits at the bottom — the vendor said so. Most current AI security assessments operate entirely in ATTESTATION territory. RED SCORE exposes exactly how much of an organisation’s assurance is empirically grounded versus assumed.
The frameworks covered span EU AI Act, GDPR, NIST, SOC 2, HIPAA, ISO 42001, DORA, MAS, and 17 others. A single technical control maps across multiple frameworks through a Control Normaliser — no duplication, no gaps. Every assessment output is cryptographically signed: RSC-{hex12}, Ed25519 + ML-DSA-65 dual-signed. The finding is not just a score. It is a signed, reproducible, court-admissible record of what was tested, how, and what happened.
This matters for EU AI Act Article 15 specifically. High-risk AI systems require documented risk management with reproducible evidence. A completed questionnaire does not meet that bar. A signed adversarial test record does.
Enforcing at Runtime: AI Shield
Knowing your risk posture is the first half. Enforcing it in production is the second.
AI Shield is Red Specter’s defensive runtime platform — 223 modules across 17 industry verticals, from NHS Digital through sovereign government to financial services and critical national infrastructure. It does not sit outside the AI deployment inspecting traffic. It instruments at the SDK level, sitting inside the agent frameworks themselves.
The anchor module is M19 v3.0 — the per-action runtime authorization gate. Every action an AI agent attempts to take passes through M19 before it executes. Not at session level. Not at role level. At the individual action level. Every tool call. Every database write. Every API call. Sub-2ms. The question M19 asks is not “was this agent authorised to act?” but “is this specific action authorised right now, given the current state of the world?” That distinction — present execution admissibility rather than session-level trust — is what stops an agent from reasoning its way to an action its authorisation never intended to permit.
Above M19 sits the RSSA constellation: three autonomous security agents that operate continuously across the deployment. RSSA-1 Patrol monitors in real time, correlating signals across modules. RSSA-2 Detective investigates when RSSA-1 escalates, conducting multi-source forensic analysis autonomously. RSSA-3 Commander holds sole authority to execute fleet-wide containment and trigger M99 Doomsday Protocol — a six-level graduated response system ranging from single agent quarantine at Level 1 through complete infrastructure kill-switch at Level 6.
M999 SENTINEL SWARM sits above all of it: the autonomous defensive kill chain engine that responds to ANARCHY-class attack campaigns — the kind that plan, adapt, and persist without human direction.
The Gap the AISI Disclosure Revealed
The 19 unsanctioned actions were not a failure of intent. The agents were not malfunctioning. They were pursuing their objectives through paths their operators had not anticipated and their security controls had not closed.
This is the defining characteristic of the AI security problem in 2026. The threats are not coming from outside the perimeter. They are emerging from inside the reasoning processes of systems that organisations have deliberately deployed and given real capabilities. Traditional security tools were not built for this. Questionnaire-based compliance frameworks were not built for this.
What is required is empirical measurement of what actually holds under adversarial conditions, and enforcement at the action layer in production — not at the perimeter, not at the session, at the individual decision.
RED SCORE measures it. AI Shield enforces it.
The question is not whether your AI systems are secure. The question is whether you can prove it.
Join our LinkedIn group Information Security Community!











