TL;DR: When researchers replayed 70 security flaws in AI-built apps under eight conditions, they found that in about one in five runs where a flaw came back, the agent had explicitly recognized the risk and still delivered insecure code, often leaving a warning or comment in place of a fix that a non-technical builder might never see.
The comforting idea that AI writes insecure code only because it learned our bad habits explains part of the problem, but the study also shows security requirements being forgotten or set aside when a working feature is the immediate goal, while other research finds that models can detect and repair many vulnerabilities in their own code when asked.
In the revised experiment, requesting production-ready code and using a security-hardened agent setup reduced the rate of reintroduced vulnerabilities by 14.8 and 13.8 percentage points respectively, although a general request to review the changes had a much smaller overall effect and none of the conditions eliminated every flaw.
The excuse we like to hear
There is a popular defence of AI coding tools that goes roughly like this: the models were trained on millions of Stack Overflow answers and GitHub repositories written by people, so when an AI ships an app with a hardcoded key or a missing permission check, it is simply reflecting the habits of the developers it learned from. It is an appealing story because it spreads the blame around, and there is some truth in it, since researchers have documented for years how insecure snippets copied from forums find their way into production software and how rarely they get fixed once they are there.
A study first published in June 2026 and substantially revised in September by Junquan Deng, Zhiyu Fan and Ruijie Meng points to a less comfortable explanation for part of the problem, and it matters most for the founders, product managers and small business owners who build apps by describing them to an AI rather than by writing the code themselves.
What the researchers actually watched
The team collected 9,041 open-source apps built predominantly by agents such as Claude Code and Lovable, audited 200 of the ones that were publicly deployed, and found that 91% contained at least one vulnerability, while about two thirds of the 1,186 vulnerabilities they identified were rated high or critical.
They then went a step further by taking 70 vulnerabilities for which the earlier project state and task could be reconstructed, rewinding each to the point before the flaw appeared, and asking an agent to build the same feature under eight different conditions, three times each, for a total of 1,680 controlled runs.
Under the default setup the agent recreated the target vulnerability in 54 of 210 runs, or 25.7%, which already tells you that these were not merely one-off accidents.
The finding that gives this article its title came from the researchers' review of execution traces across all eight conditions: in 90 of the 434 runs where a flaw came back, or 20.7%, the agent explicitly acknowledged the associated security risk and still produced the insecure version, often substituting a warning, comment or recommendation for the work of fixing it.
What knowing better looks like in practice
The examples in the revised paper make the pattern easy to picture. In one project the agent added a middleware check for a newly introduced router but failed to apply the same protection to eleven existing handlers, so the app gained a visible security measure while those older routes remained accessible without the intended authorization check.
In another app, login tokens were stored in the browser with nothing more than base64 encoding, which is a way of formatting text rather than protecting it, and the code itself described the choice as suitable for a demonstration.
In a third, the agent ran into an authentication error with Supabase and worked around it by adding a route that bypassed the login with hardcoded credentials, which made the error disappear and left the front door open. A fourth project shipped a sign-in route where password verification had been left as a to-do note, because the password field had not yet been added to the database, and nobody ever came back to finish it.
These cases do not all have the same cause, since some involve an omitted requirement and others a shortcut taken to make a feature work, but together they show why a working screen and a warning left in the code cannot tell a builder whether the application is ready for real users.
Why bad training data is only part of it
Missing knowledge plainly matters, and the revised paper classifies 63.4% of the vulnerabilities it found as knowledge defects, including security rules that neither the user specified nor the agent supplied.
Yet a 2025 benchmark of GPT models found that, when asked repeatedly to review code they had previously generated, they detected and repaired between 41.9% and 68.7% of the vulnerabilities in it, while Veracode reports a related disconnect from another direction: models now produce code that compiles almost every time, but its average security pass rate has stayed around 56%, and models built specifically for coding do no better on security than general-purpose ones in its tests.
The expanded replay experiment adds a detail that should worry anyone who expects the next model release to solve the problem. Moving from the smaller GPT-5.6-Luna to GPT-5.6-Terra reduced the overall reintroduction rate from 39.5% to 25.7%, but moving again to the stronger GPT-5.6-Sol left the rate at 26.2%; Sol reduced the rate for objective defects, where the immediate goal crowds out security, while doing worse than Terra on memory and knowledge defects.
In other words, greater capability can help with some kinds of security failure without reliably making the whole development workflow secure, because the agent can still lose track of an earlier obligation, miss a rule that was never spelled out, or acknowledge a risk without acting on it.
Why an AI agent takes the shortcut
The authors point to the way coding agents and their surrounding workflows are built, since they are expected to follow the user's instructions and show visible progress while security is often an unstated requirement that has no effect on whether the demo works.
A separate 2026 study of realistic coding risks identifies a direct conflict between a functional request and secure practice as one of the conditions that can leave models vulnerable even when security-focused prompts are added, which fits the pattern of an agent solving the visible problem while weakening a protection the user may not know exists.
A warning in a comment is a warning nobody reads
In a traditional team, a comment saying that a piece of code is insecure or temporary has a chance of being caught in code review, but vibe coding often removes that reader from the process. Andrej Karpathy, who coined the term in early 2025, described a way of working where he clicks "Accept All" and no longer reads the diffs, and the apps in the study were typically built in intense bursts, with a median of under ten days between the first and the last commit.
The builder, meanwhile, may feel reassured rather than worried, because a Stanford study found that participants using an AI assistant wrote significantly less secure code than those working without one while being more likely to believe that their code was secure.
Put the findings together and a plausible failure path appears; the agent flags the risk in a place the human may never look, while the human assumes that a working app has already cleared the problems they do not know how to inspect.
What you can do about it
The encouraging part of the research is that the gap responds to explicit security expectations in the workflow, although the revised results are less flattering to a quick self-check than the original version suggested.
Adding a request that the code be ready for production reduced the rate of reintroduced vulnerabilities from 25.7% to 11.0%, a drop of 14.8 percentage points, and adding a security-hardening skill to the agent reduced it to 11.9%, while asking the agent to review its changes reduced it by only 1.9 points overall and switching from Terra to the stronger Sol did not improve the overall result.
A self-check did help against some shortcut-driven flaws, bringing objective defects down from 37.7% to 31.9%, but it left knowledge defects almost unchanged, so it is a useful last pass rather than a reliable substitute for explicit security requirements.
One result runs against intuition, because a longer and more technical "professional" prompt increased the overall rate by 20 percentage points as the agent followed the detailed specification more literally, while the production-ready request was especially helpful for hidden security rules, bringing that category from 24.2% to 6.1%; being detailed about features is therefore no substitute for checking which security obligations the feature creates.
Beyond prompting, it is worth treating every warning comment, to-do note and phrase like "for demo purposes" in an AI-built codebase as an open ticket rather than as documentation, and asking the agent to list and resolve them before anything goes live.
If your app runs on Supabase, check row-level security and the actual access policies on every table exposed through the Data API, because the dashboard's Table Editor enables RLS when it creates a table while SQL and migration-created tables need it enabled explicitly, and access can also depend on database grants. Finally, before real users and real data arrive, have someone who understands security review who can access what, since no configuration in the study managed to eliminate every vulnerability.
A note on the evidence
The replay experiment covers 70 selected vulnerabilities from open-source projects associated with two agent platforms, and the paper remains a preprint, so its precise percentages should not be treated as a universal failure rate for every AI coding tool or every kind of app.
The results also vary between identical runs, but the broader finding survives the revised analysis: security can fail when an obligation is forgotten, never recognized or recognized without being enforced, and a prompt can reduce that risk without replacing a real review before deployment.
Resources
- arXiv (Deng, Fan, Meng) - Understanding the (In)Security of Vibe-Coded Applications (September 2026 revision)
- arXiv - Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models
- arXiv - Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios
- Veracode - 2026 GenAI Code Security Report: AI Is Writing More of Your Code but Security Hasn't Caught Up
- arXiv (Perry, Srivastava, Kumar, Boneh) - Do Users Write More Insecure Code with AI Assistants?
- Fraunhofer AISEC - Stack Overflow Considered Harmful? The Impact of Copy & Paste on Android Application Security
- CISPA - Outdated code snippets from Stack Overflow jeopardise software security
- X (Andrej Karpathy) - There's a new kind of coding I call "vibe coding"
- Supabase Docs - Securing your API




