Skip to main content
1PasswordLaunch

AI struggles to patch vulns without adult supervision

What happened

1Password's Off-by-1 Labs released a research paper and a patch evaluation harness named FLAWED to help organizations evaluate the effectiveness of their security fixes.

Source

Article excerpt

AI and ML Left alone, autonomous fixes often fail to fully remediate flaws AI models may not be that good at fixing security flaws. Researchers at 1Password's Off-by-1 Labs analyzed security patches generated by two frontier models - ChatGPT 5.5 at "medium" effort and Claude Opus 4.8 at "high" effort - and found that autonomous patches cleanly fixed vulnerabilities only about a quarter of the time, while most of the remainder failed to fully remediate the flaw or introduced other problems. Keith Hoodlet, director of security research at 1Password, argues in a blog post that the results show LLM-driven security remediation still needs human review. "Across six recently disclosed CVEs, we produced 6,080 patches using two frontier, cyber-capable reasoning models," Hoodlet said. "The average success rate for generating a patch that fully resolved the vulnerability (without materially changing application behavior) was just 26.0 percent." Of the AI-generated patches, 20.1 percent fixed the original issue but altered application behavior (eg, changing "allow list" logic to "deny list" logic). Some 2.3 percent of the patches fixed the issue while introducing new security issues. 49.3 percent of the patches failed to fix at least one existing exploit path. And 2.2 percent both failed to fix the vulnerability while introducing a new exploit path. And among the patches in the first...

Keep reading with a free account

The rest of this article, and every signal for 1Password, is in your free account.

Extracted from this sentence

The authors have released a patch evaluation harness under the name FLAWED that organizations can use to evaluate the effectiveness of their security fixes.

Extracted by Autobound

From the Signal API record
Event
Launch

What this signalsA launch often needs new go-to-market and support spend.

Product
FLAWED

The full record

From the Signal API record

Topics and mentions

Product tags

  • security
  • general technology

Extraction

Confidence
90%
Detected
Aug 6, 2026
signal_type
news
signal_subtype
launches

Use this data

Get every launch signal for 1Password and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at 1Password this week?”

  2. Send it to your own tools

    The Signal API returns launch signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full news record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/5b053da1-e6c5-b8a6-1904-bab36fdc1f4d returns this record as JSON. POST /v1/companies/enrich returns every signal for 1password.com.

{
  "signal_id": "5b053da1-e6c5-b8a6-1904-bab36fdc1f4d",
  "signal_type": "news",
  "signal_subtype": "launches",
  "detected_at": "2026-08-06T19:04:30+00:00",
  "company": {
    "name": "1Password",
    "domain": "1password.com"
  },
  "data": {
    "url": "https://www.theregister.com/ai-and-ml/2026/08/06/ai-struggles-to-patch-vulns-without-adult-supervision/5284319",
    "title": "AI struggles to patch vulns without adult supervision",
    "excerpt": "AI and ML Left alone, autonomous fixes often fail to fully remediate flaws AI models may not be that good at fixing security flaws. Researchers at 1Password's Off-by-1 Labs analyzed security patches generated by two frontier models - ChatGPT 5.5 at \"medium\" effort and Claude Opus 4.8 at \"high\" effort - and found that autonomous patches cleanly fixed vulnerabilities only about a quarter of the time, while most of the remainder failed to fully remediate the flaw or introduced other problems. Keith Hoodlet, director of security research at 1Password, argues in a blog post that the results show LLM-driven security remediation still needs human review. \"Across six recently disclosed CVEs, we produced 6,080 patches using two frontier, cyber-capable reasoning models,\" Hoodlet said. \"The average success rate for generating a patch that fully resolved the vulnerability (without materially changing application behavior) was just 26.0 percent.\" Of the AI-generated patches, 20.1 percent fixed the original issue but altered application behavior (eg, changing \"allow list\" logic to \"deny list\" logic). Some 2.3 percent of the patches fixed the issue while introducing new security issues. 49.3 percent of the patches failed to fix at least one existing exploit path. And 2.2 percent both failed to fix the vulnerability while introducing a new exploit path. And among the patches in the first two...",
    "product": "FLAWED",
    "summary": "1Password's Off-by-1 Labs released a research paper and a patch evaluation harness named FLAWED to help organizations evaluate the effectiveness of their security fixes.",
    "planning": false,
    "image_url": "https://image.theregister.com/5255449.jpg?imageId=5255449&x=0&y=0&cropw=100&croph=100&panox=0&panoy=0&panow=100&panoh=100&width=1200&height=683",
    "confidence": 0.9,
    "product_data": {
      "name": "FLAWED",
      "full_text": "a patch evaluation harness under the name FLAWED",
      "fuzzy_match": false
    },
    "product_tags": [
      "security",
      "general_technology"
    ],
    "published_at": "2026-08-06T19:04:30Z",
    "article_sentence": "The authors have released a patch evaluation harness under the name FLAWED that organizations can use to evaluate the effectiveness of their security fixes."
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.