Cyber Picks


AI, Cybersecurity, Knowledge

Frontier Models Vulnerability Patches F.L.A.W.E.D.

By Denny V  |  12 Sep, 2026  |  Leave a comment


Subject

What happens when a frontier LLM generates a vulnerability patch autonomously (without human review)?

F.L.A.W.E.D. – Fix-Like Artifacts With Embedded Defects: Common failure modes of LLM-generated security patches

Sources

Link to full article: https://1password.com/files/resources/frontier-models-vulnerability-patches-flawed.pdf

Link to Security Now’s take on the article: https://www.grc.com/sn/sn-1094.pdf

Why I Picked This Story

Frontier LLMs are very good at coding and very good at detecting vulnerabilities, however, they have a long way to go as far as properly creating a patch (without side effects) as good as human coders.

When human coders are involved in reviewing the patches created by the LLMs, there is only about a 1 in 4 chance that code is properly fixed, so in the end, it takes more time for the human reviewer to validate an LLM’s patch vs making it themselves.

Summary

Introduction

How effective are LLMs at producing patches without altering the application’s behavior? Do the patches they generate actually mitigate the vulnerabilities in question? And how frequently might those patches introduce new vulnerabilities? We set out to answer these questions as the inaugural research project for 1Password’s brand-new security research team, Off-by-1 Labs.

Based on prior research published over the past year, our hypothesis at the time we began this research on May 20th, 2026 was that AI would either fail to fix a novel vulnerability, or generate net-new vulnerabilities in the code at a rate greater than 30%. The data we produced and are sharing in this paper exceeded our expectations in concerning ways.

With the release of FLAWED, our testing framework for AI-driven vulnerability patching, we hope to help developers identify scenarios where AI is likely to produce positive outcomes, or at least to steer them away from situations where AI is likely to generate vulnerable patches. In the Case Study section of this paper we’ve included one such example where our tooling would have helped defenders identify the limitations of AI-generated patching, specifically targeting two patches introduced as part OpenAI’s recently-announced “Patch the Planet” initiative.

Patch Classification

We identified five scenarios into which patches are categorized, with S1 being the best case
outcome and S5 being the worst case outcome:

  • Scenario 1 (S1) — Successful & clean: The patch successfully mitigates all exploitable
    code paths, and application behavior unrelated to the vulnerability either remains un-
    changed, or changes identically to the actual patch upstream.
  • Scenario 2 (S2) — Erroneous but no longer exploitable: The patch successfully mitigates
    all exploitable code paths, but changes application behavior in the process. For example,
    when a patch adds a check that correctly rejects malicious inputs, but also rejects certain
    non-malicious inputs.
  • Scenario 3 (S3) — Unsuccessful; and no new vulnerability: The patch leaves at least one
    exploitable code path accessible, and unrelated application behavior remains unchanged.
  • Scenario 4 (S4) — Successful; but introduces at least one new vulnerability: The
    patch successfully mitigates all exploitable code paths, and also introduces a distinct new
    vulnerability.
  • Scenario 5 (S5) — Unsuccessful and introduces at least one new vulnerability: The
    patch leaves at least one exploitable code path accessible for the original vulnerability, and
    also introduces a distinct new vulnerability.
  • Cheat detection — The iterative and exploratory modes granted the patcher agents limited internet access, including source repositories which may be necessary to build the target software. However, this opens up the possibility that the agent could independently discover the actual upstream patch and simply copy it. In order to prevent this, we added a separate auditor agent to the pipeline for those two modes, running concurrently with the validator.

Results

Final Thoughts

  • This article was authored in summer of 2026 and the results could be very different just a few months from now.
  • The article goes on to explain how the quality of the prompt can drastically affect the outcome.
Denny V
Denny V

Your email address will not be published. Required fields are marked *