Skip to content
← Back to blog
Engineering·October 1, 2026·6 min read

AI models now pass SQL injection tests 83% of the time and cross-site scripting tests 15% of the time. Veracode's 2026 data says review for the second one.

Veracode's July 2026 report: AI code security is stuck at a 56% pass rate, with 83% on SQL injection but 15% on XSS and 12% on log injection.

Veracode's 2026 GenAI Code Security Report, published July 28, 2026, found that the average security pass rate across AI models is 56%, "barely changed from 55% in the first report." That headline is easy to read as "AI code is about half safe" and move on. The more useful reading is inside the average. the models Veracode tested pass SQL injection tests 83% of the time and cryptographic-algorithm tests 87% of the time, and they pass cross-site scripting tests 15% of the time and log injection tests 12% of the time. If your engineers review AI-written code for the flaws everyone already knows to look for, the data says you are reviewing the part that mostly works.

What Veracode measured

Veracode is an application security vendor, and it publishes this report as a recurring benchmark of security in AI-generated code. The 2026 edition tested 11 new models across 80 coding tasks, and Veracode says the full series now covers more than 100 models tested over four years. A task counts as failed when the generated code introduces a security vulnerability, and the report's blog summary puts the result plainly: roughly 44% of AI code generation tasks introduced a risky vulnerability. Two caveats belong next to that. This is a vendor's own benchmark, not an independent audit, and 80 tasks is a fixed test set, not a sample of your codebase. Treat the numbers as a map of where models are weak, not a forecast of your defect rate.

Why a flat 56% matters more than a low one

A pass rate that started at 55% and sits at 56% after four years of new models is the finding. Veracode's report says security has stayed flat while the amount of AI-generated code entering pipelines has grown, which turns a benchmark curiosity into a volume problem. The company also reports that the best model in its Summer 2026 dataset, GPT-5.5, passes 68% of tasks, so even the leader fails nearly one in three, and six of the eleven models sit between 50% and 53%. Model choice moves you a little. It does not move you out of the problem.

The type of model barely helps either. Veracode reports code-specialised models averaging 51% and general-purpose models 52%. Reasoning models average 56% against 51% for non-reasoning ones, a real but small edge, and large models average 53% against 51% for medium and small. If a vendor tells you their newer model made security a solved concern, the publisher of the longest-running benchmark on the subject disagrees.

The split that changes how you review

Here is what the per-vulnerability numbers say, as reported by Veracode in July 2026:

Vulnerability classAverage pass rate across models
Cryptographic algorithms87%
SQL injection83%
Cross-site scripting (XSS)15%
Log injection12%

SQL injection and weak cryptography are the flaws every developer has been warned about for twenty years, so they are heavily represented in what models learned from. Cross-site scripting and log injection are the ones that require the model to know that a value came from a user and has to be sanitised on the way out or into a log line. That is a data-flow judgement, and it is exactly what a model writing one function at a time does not see. The review checklist most teams carry (parameterised queries, no home-made crypto) covers the two classes where models already do fine. It skips the two where they mostly fail.

Language matters too. Veracode's report calls Java the language with the clearest improvement trend and also the one that is "last by a wide margin," with a mean security pass rate of only 30%. If your product is a Java backend with AI-assisted changes, your baseline is worse than the 56% average, not better.

A worked example

Picture a 30-person Series A company whose product takes free-text input from customers, stores it, and renders it back in a dashboard, plus an audit log line for every action. An engineer uses an AI assistant to add a "notes" feature. The generated code uses parameterised queries, and the reviewer, checking for injection the way the team always has, nods it through. The note text is rendered into the page unescaped and written into the log line verbatim. Neither problem is exotic. Both sit in the two classes where Veracode's tested models fail most often, and both passed a review aimed at the two classes where they don't. Nothing here is a reason to stop using the assistant. It is a reason to point the reviewer somewhere different.

What to change on Monday

Three changes cost little. First, add output encoding and log sanitisation to the pull request template as explicit questions, so the reviewer is asked about the classes that fail, not the ones that pass. Second, turn on static analysis rules for XSS and log injection in CI and make them blocking, since a human skimming a diff is the weakest control for a data-flow bug. Third, track which changes were AI-assisted long enough to see whether your own failure pattern matches the benchmark's. Our earlier pieces on why AI-assisted review is now the bottleneck and on AI code review quality against production incidents cover the throughput side. This one is about what the reviewer should be looking for once the diff is in front of them. For the attack surface on the model side, see our prompt injection breakdown.

Where this doesn't apply, and where help fits

If your team is four engineers who all read every line, a written checklist and a CI rule are enough, and bringing in outside help for this would be waste. Extra hands earn their cost when review capacity is the constraint: a senior reviewer who knows these classes is a scarce hire, and the queue of AI-assisted diffs grows faster than they can read. That is the case where a dedicated security-minded engineer embedded in your team, or a scoped hardening pass on the features that take user input, beats another quarter of hoping. Our Silicon Valley team builds these under build your team, and you can tell us where your review queue is stuck and get a plain answer on whether it needs people or just a better checklist.

Sources

Frequently asked questions.

Veracode's 2026 GenAI Code Security Report, published July 28, 2026, found an average security pass rate of 56% across models, barely changed from 55% in its first report. The 2026 edition tested 11 new models across 80 coding tasks, and the top model in its Summer 2026 dataset still failed nearly one in three tasks.

In Veracode's July 2026 report, models averaged a 15% pass rate on cross-site scripting and 12% on log injection, against 83% on SQL injection and 87% on cryptographic algorithms. The weak classes are the ones that depend on tracking untrusted input through the code, not on avoiding a well-known bad pattern.

Only slightly. Veracode's 2026 report shows reasoning models averaging a 56% pass rate against 51% for non-reasoning models, large models at 53% against 51% for medium and small, and code-specialised models at 51% against 52% for general-purpose ones. The overall average has moved from 55% to 56% across the report's four-year series.

Aim the reviewer at the classes where models fail most. Veracode's July 2026 data shows near-zero pass rates for cross-site scripting and log injection, so pull request checklists should ask about output encoding and log sanitisation, and CI should run static analysis rules for both as blocking checks instead of relying on a reader to spot them in a diff.

It is a map of model weaknesses, not a forecast. Veracode is a security vendor running its own fixed set of 80 tasks, so the results are not an independent audit or a sample of any one company's code. They are useful for deciding where review effort should go, and your own tracking of AI-assisted changes will show whether the pattern holds for you.