Skip to main content

Can AI-Generated Code Be Trusted? Research on Security Vulnerabilities in Code Written by AI Models

Empirical audits reveal that up to 40% of LLM-generated code contains critical vulnerabilities, while developer overconfidence and package hallucinations expose enterprise software supply chains.

By Vodnala Akshith
Published: Oct 05, 2026
7 mins read
👁️ 19 Unique Views
Can AI-Generated Code Be Trusted? Research on Security Vulnerabilities in Code Written by AI Models
The scale of inference: Optimized for multimodal workloads.
Premium Insight

Why It Matters

AI code generators accelerate development velocity, but without rigorous automated verification, they equally accelerate the accumulation of catastrophic vulnerabilities. Syntactic fluency must never be equated with security. Enterprise engineering teams must treat all AI-generated code as untrusted input, enforce automated SAST/DAST verification within CI/CD loops, and implement strict dependency controls to protect against package hallucinations.

The rapid proliferation of large language model (LLM) coding assistants—such as GitHub Copilot, Claude 3.5 Sonnet, OpenAI Codex, and CodeLlama—has transformed commercial software engineering. By accelerating boilerplate generation, automating API scaffolding, and synthesizing multi-file modules from natural language prompts, neural code generation has delivered measurable developer velocity gains. However, this productivity boom conceals a dangerous structural hazard: Large Language Models optimize strictly for token probability, predicting tokens that mirror public software repositories rather than verifying semantic soundness or operational security.

Because models generate syntactically pristine, beautifully formatted code that compiles without warnings, engineers routinely fall prey to an illusion of competence. In human software engineering, syntax errors and code smells frequently correlate with cognitive lapses and logical flaws; developers use these cues to signal when a module warrants intense inspection. With neural code synthesis, this diagnostic link is severed. Models effortlessly output clean, idiomatic routines that conceal critical vulnerabilities, such as off-by-one memory corruptions, improper authentication bypasses, and unescaped database interpolations.

"Asleep at the Keyboard": Empirical Vulnerability Rates

The foundational empirical benchmark assessing LLM code security was conducted by researchers at NYU Tandon (Pearce et al., IEEE Symposium on Security and Privacy 2022), titled "Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions". The authors methodically tested Copilot across 89 distinct programming scenarios targeting the top-25 Common Weakness Enumeration (CWE) categories, generating nearly 1,700 unique programs across Python, C, and JavaScript.

The findings were striking: across all evaluated scenarios, 39.33% of the generated programs were demonstrably vulnerable to high-severity security exploits. In low-level languages like C, Copilot routinely emitted buffer overflows (CWE-119/120) and pointer arithmetic flaws. In web-oriented stacks, Copilot frequently generated textbook SQL injection vectors (CWE-89) by interpolating raw user input directly into SQL strings, alongside Cross-Site Scripting (CWE-79) and Path Traversal flaws (CWE-22). The model did not synthesize novel attack surfaces; rather, it functioned as a faithful statistical mirror of the historically vulnerable patterns that saturate public open-source codebases.

The Overconfidence Trap: Human-in-the-Loop Degradation

Advocates of AI-assisted engineering often contend that as long as a human programmer remains "in the loop," LLM vulnerabilities pose minimal risk because qualified developers will review and sanitize every pull request. However, a landmark human-subjects investigation by Stanford University (Perry et al., ACM CCS 2023), "Do Users Write More Insecure Code with AI Assistants?", demonstrated that human oversight actually deteriorates in the presence of AI assistants. Evaluating 47 developers across five security-critical tasks including symmetric encryption, digital signatures, user authentication, and database querying, the researchers uncovered two alarming behavioral trends:

  • Degraded Security Outcomes: Participants with access to AI assistants produced statistically significantly more insecure code than unassisted control developers. In C tasks, AI users routinely omitted buffer length checks; in Python, they defaulted to broken cryptographic ciphers (e.g., DES or ECB-mode AES).

  • The Overconfidence Paradox: Despite writing substantially more vulnerable code, participants in the AI-assisted cohort expressed statistically higher confidence that their implementations were secure compared to unassisted peers. The aesthetic fluency of AI completions induced cognitive complacency, leading developers to rubber-stamp lethal logic flaws.

Supply-Chain Hallucinations: The "Slopsquatting" Threat Vector

Beyond localized code-level flaws, LLMs introduce systemic vulnerabilities into the broader software supply chain through package hallucination. When prompted to solve complex, niche, or newly emerging software challenges, models frequently hallucinate plausible-sounding third-party libraries (e.g., Python packages on PyPI or JavaScript modules on npm) that do not actually exist in public registries.

Empirical research by cybersecurity researcher Bar Lanyado (Vulcan Cyber) and subsequent university audits proved that attackers can systematically harvest these hallucinated package names from model queries. By registering these phantom package names on PyPI or npm with embedded trojans—a vector dubbed "Slopsquatting"—adversaries achieve remote code execution (RCE) inside enterprise developer environments and CI/CD pipelines whenever developers or autonomous coding agents execute automated installation commands like pip install.

Root Causes: Why Large Language Models Generate Vulnerable Code

Understanding why neural models generate insecure code requires deconstructing their operational primitives. First, corpus contamination is unavoidable: public GitHub repositories and StackOverflow archives are filled with legacy, unmaintained, or demonstrably vulnerable code written before modern defensive paradigms were established. Second, models lack whole-repository threat modeling; an LLM generates completions within a restricted token context window without understanding global authentication middleware, network isolation boundaries, or sanitization layers.

Engineering Defenses: SAST-in-the-Loop and Constrained Decoding

Securing AI-assisted software development requires replacing passive human review with deterministic systems-level guardrails. Leading engineering organizations are implementing closed-loop agent architectures:

  • Dual-Agent Auditor Loops: Coupling code-generation models with dedicated security auditor agents running Static Application Security Testing (SAST) engines (e.g., Semgrep, CodeQL, Bandit) before code is accepted.

  • Constrained Grammar Decoding: Enforcing strict AST schemas that forbid dynamic string interpolation in database drivers, requiring parameterized query builders at the token decoding stage.

  • Strict Dependency Sandboxing: Rejecting unverified package imports and restricting agent build environments with strict network egress policies and private package mirror registries.

Found this analysis insightful?

Share with colleagues, engineers, and your network.

Link copied to clipboard!