
The reasoning that introduces a flaw is often the reasoning asked to catch it.
In March 2026, an attacker backdoored the GitHub Action for Trivy, a widely used open source security scanner. The compromised action ran inside the CI/CD pipeline of LiteLLM, the gateway that routes model requests for CrewAI, DSPy, Microsoft GraphRAG, and a long list of other agent frameworks. There it harvested the credentials LiteLLM uses to publish to PyPI. The attacker pushed two malicious releases straight to the package index, and although they were live for well under an hour before PyPI quarantined them, researchers still put the attack surface reach at more than 2,000 organizations. A security scanner became the delivery vehicle for a compromise of the exact plumbing agentic AI systems depend on.
It’s why OWASP’s GenAI Security Project ranked prompt injection the top risk in its Top Ten for the third consecutive year, this time backed by a database of roughly 10,000 real-world incidents. As Steve Wilson, the project’s co-chair, put it, this year’s update is “grounded in much more than expert opinion.” The architectural root cause hasn’t moved in three years: a large language model processes instructions and data through the same channel, and it cannot reliably tell your intent from an attacker’s.
The vulnerabilities don’t start at runtime, either. Stanford researchers demonstrated back in 2022 that developers using an AI coding assistant wrote measurably less secure code than developers working without one, and the AI-assisted group was more confident their code was secure, not less.
That study is old enough now that the obvious objection is a fair one: the models have improved enormously since. The trouble is that the newer measurements point in the same direction. Veracode’s evaluation of more than 100 models across four languages found AI-generated code carries 2.74 times the vulnerabilities of human-written code and fails secure-coding benchmarks 45% of the time. It’s 2026 update found that newer, larger models scored no better than smaller ones. Apiiro, studying real Fortune 50 codebases, found 322% more privilege escalation paths and 40% more secrets exposure in AI-assisted repositories. Capability went up. The gap between confidence and correctness didn’t close, and the code those developers ship now moves through a review process built for a slower, more skeptical era of software development.
Package hallucination is the clearest evidence of how that compounding plays out. Researchers from the University of Texas at San Antonio, Virginia Tech, and the University of Oklahoma tested 16 code-generating models across 576,000 samples and found that 19.7% referenced a package that doesn’t exist, more than 205,000 unique fictional names in total. Forty-three percent of those hallucinated names showed up every time the same prompt ran again, which is precisely why security researcher Seth Larson coined the term slopsquatting for what happens next: an attacker doesn’t need to guess which fake package name an AI assistant will suggest. They observe the pattern, register the name on npm or PyPI, and wait for a developer or an autonomous agent to install it.
Black Duck’s own research puts a number on how much of this is now baked into the default workflow. The company’s 2026 State of AI-Powered Software Development report, conducted with research firm UserEvidence, found that 97% of organizations are now using AI coding assistants in production software. Adoption is high enough that prohibition has stopped functioning as a control: among the minority of companies that explicitly ban these tools, 76% acknowledge their developers use them anyway. The codebases show the consequence. Black Duck’s 2026 Open Source Security and Risk Analysis report, an audit of 947 commercial codebases rather than a survey, found that mean vulnerabilities per codebase more than doubled, rising 107% to an average of 581, the sharpest single-year jump in the report’s history. Put the survey data and the audit data next to each other, and the picture is consistent: the tools are close to universal; the failure modes are documented and repeatable, and the governance to catch either hasn’t arrived at the same pace.
None of this makes AI coding assistants the problem. I’ve spent 30 years in cybersecurity, compliance, and application security, and I’ve watched adoption outrun governance before for cloud, mobile and open source, just to name a few. What’s different this time is that human review scales linearly and AI-assisted code generation doesn’t. Pull request volume on AI-heavy teams keeps climbing while reviewer capacity stays flat, and that gap is what turns a manageable backlog into an unmanageable one within a couple of quarters. The answer to that isn’t more AI pointed at the same code. It’s making sure the controls you depend on fail differently than the thing they’re checking. Uncorrelated failure modes are the entire reason defense in depth works, and they are the first thing an AI-first security posture quietly gives up.
Three things AI still doesn’t bring to a line of code
An AI coding assistant doesn’t carry your business context. It doesn’t know what your application does in production, what data it touches, or what a breach costs you in regulatory exposure. It generates code that satisfies your prompt, with no consideration for your threat model.
It also has no architectural memory. It doesn’t know your architecture decision records or the incidents your team has already survived. Every session starts from zero, and if you don’t feed it that context deliberately, it will violate decisions your team made for reasons the model never saw.
And it isn’t optimizing for security. It’s optimizing for plausible, useful completions which is the same design property behind the insecure code, the hallucinated packages, and the inability to separate trusted instructions from attacker-controlled text. When that same system moves from suggesting code to autonomously committing it, deploying it, and managing the pipeline around it, a compromised or misdirected agent stops producing a bad pull request. It produces a production incident with no human in the loop to catch it.
Why AI reviewing AI isn’t enough on its own
Using the same class of model to check code it may have helped write creates a self-referential loop. That isn’t a claim that the second pass catches nothing. It can and does catch plenty, and a security-scoped review is worth running. The problem is correlation. Models trained on overlapping data share blind spots, so their errors cluster in the same places, and a second opinion drawn from the same distribution is not an independent one. Organizations that treat an AI security review as equivalent to an independent one are trading a known gap for one that’s harder to see.
The fix has to include methods with different failure modes, not just more models pointed at the same code. Static analysis is deterministic and repeatable in a way AI reasoning by design isn’t. Software composition analysis flags the dependency and slopsquatting risk an LLM has no reason to catch in its own output. And a person still has to read the highest-stakes diffs, because no scanner, AI or otherwise, has replaced a security engineer evaluating the trust boundaries of an agentic workflow.
Seven layers, not one scanner
The defense for this doesn’t require abandoning AI-assisted development. It requires treating it the way security teams have always treated every other high-velocity risk: with multiple layers rather than a single control.
- Ground the model in your standards, your ADRs, and your accepted risks before it writes a line.
- Threat-model twice, once for the code AI writes and once for the AI system at runtime.
- Build a test loop where every fix generates a permanent regression test.
- Run a second AI pass scoped only to security, on changed code, not the whole repo (necessary but never counted as the independent review).
- Gate SAST, SCA, and DAST in CI as required steps, not optional ones.
- Keep a human reviewing the highest-stakes diffs: auth logic, payment flows, and agentic permissions.
- Monitor at runtime, because some of this will get through anyway.
I’ve applied all seven of these to my own work. The application I built with AI assistance carries more than 7,000 automated tests across roughly 270 files, and an AI security review ran at every stage (before the code was written, after it was written, and again on every pull request diff). It still returned more than 200 findings on traditional AppSec scans including DAST, SAST, and SCA. The security scans identified three primary vulnerability categories: (1) Dependency vulnerabilities in xlsx parsing; (2) Web application security gaps addressed through WAF deployment (ModSecurity + OWASP CRS in detection mode with 11 application-specific exclusions); and (3) Infrastructure hardening findings that led to container security baselines enforced across all six Docker services (dropped capabilities, read-only filesystems, no privilege escalation). None of these are arguments against using AI to write code. It’s the argument for never treating AI-generated code as a substitute for either good software engineering design, solid AI development techniques and processes, and the scanning discipline that existed before AI showed up.
The organizations getting this right aren’t the ones with the best prompts. They’re the ones who accepted, months ago, that a 97% adoption rate without a matching governance rate is just exposure with better tooling attached.
Join our LinkedIn group Information Security Community!











