Skip to main content
WizCustomer feedback

A commenter praised Wiz's new Cyber Model Arena for its practical approach of tying AI agent success rates to the cost per solve, which helps users determine which setups are practical for...

What happened

A commenter praised Wiz's new Cyber Model Arena for its practical approach of tying AI agent success rates to the cost per solve, which helps users determine which setups are practical for repeated use.

Source

RedditSep 13, 2026By u/MushroomRight283

r/aiagents

Security benchmarks are starting to expose how much the harness matters

upvotes
39
comments
11

Extracted from these lines

  • [comment u/InternNo8266] This is a better way to look at security agents than raw benchmark scores. Wiz tying success rate to cost per solve makes it easier to see which setups are practical to run repeatedly

reddit.com/r/aiagents/comments/1wfmkvb/security_benchmarks_are_starti...Open the source

Comments on the post

5 of 11 comments
  • “How are they normalizing the harness side here? If Wiz is running the same task set across ADK ReAct and Claude Code with the model held constant, that makes the gap interesting compared to a normal leaderboard. Would be good to know what counts as a solve in the Arena too”

    u/Character_Event_45375 points · Sep 14, 2026View

  • “This is a better way to look at security agents than raw benchmark scores. Wiz tying success rate to cost per solve makes it easier to see which setups are practical to run repeatedly”

    u/InternNo82661 points · Sep 14, 2026View

  • “good to see this getting measured rather than argued about. harness design ends up mattering as much as model choice and there is very little written on it.”

    u/_Ojin1 points · Sep 14, 2026View

  • “This matches what I see. Same model, two harnesses, very different injection success rates, because the harness decides what the model can actually reach. The two things that move the number for me are whether tool output gets fed back as plain text into the same context as instructions, and whether the harness re-confirms a write after a tool result changes the plan mid-run. A benchmark that”

    u/Available_Teaching831 points · Sep 14, 2026View

  • “The harness is significantly more important than the model. Not sure how people don't understand this...”

    u/seventyfivepupmstr1 points · Sep 20, 2026View

Extracted by Autobound

From the Signal API record
Signal
Customer feedback

What this signalsUser posts often show product pain before it reaches reviews or churn.

Subreddit
r/aiagents

The full record

From the Signal API record

Numbers

Mentions
2

Details

Timing
In progress
Category
Features
Link URL
wiz.io
Virality
Somewhat high
Post kind
Link
Prominence
Core
Company's role
Vendor

Topics and mentions

Topics

  • security
  • benchmarking
  • cost analysis
  • ai

Flair

  • News

Products named

  • Cyber Model Arena

Extraction

Sentiment
Positive
Detected
Sep 13, 2026
signal_type
reddit-company
signal_subtype
customerFeedback

Use this data

Get every Reddit signal for Wiz and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at Wiz this week?”

  2. Send it to your own tools

    The Signal API returns Reddit signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full reddit-company record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/81877235-f53c-5967-a272-ed1a382ff46d returns this record as JSON. POST /v1/companies/enrich returns every signal for wiz.io.

{
  "signal_id": "81877235-f53c-5967-a272-ed1a382ff46d",
  "signal_type": "reddit-company",
  "signal_subtype": "customerFeedback",
  "detected_at": "2026-09-13T23:05:19+00:00",
  "company": {
    "name": "Wiz",
    "domain": "wiz.io"
  },
  "data": {
    "nsfw": false,
    "stage": "none",
    "awards": 0,
    "timing": "in_progress",
    "topics": [
      "ai",
      "security",
      "benchmarking",
      "cost analysis"
    ],
    "post_id": "1wfmkvb",
    "summary": "A commenter praised Wiz's new Cyber Model Arena for its practical approach of tying AI agent success rates to the cost per solve, which helps users determine which setups are practical for repeated use.",
    "category": "features",
    "comments": [
      {
        "url": "https://www.reddit.com/r/aiagents/comments/1wfmkvb/comment/p9nvdz8/",
        "depth": 0,
        "score": 5,
        "author": "Character_Event_4537",
        "excerpt": "How are they normalizing the harness side here? If Wiz is running the same task set across ADK ReAct and Claude Code with the model held constant, that makes the gap interesting compared to a normal leaderboard. Would be good to know what counts as a solve in the Arena too",
        "posted_at": "2026-09-14T00:48:14.000Z",
        "author_url": "https://www.reddit.com/user/Character_Event_4537/"
      },
      {
        "url": "https://www.reddit.com/r/aiagents/comments/1wfmkvb/comment/p9nsxxk/",
        "depth": 0,
        "score": 1,
        "author": "InternNo8266",
        "excerpt": "This is a better way to look at security agents than raw benchmark scores. Wiz tying success rate to cost per solve makes it easier to see which setups are practical to run repeatedly",
        "posted_at": "2026-09-14T00:35:10.000Z",
        "author_url": "https://www.reddit.com/user/InternNo8266/"
      },
      {
        "url": "https://www.reddit.com/r/aiagents/comments/1wfmkvb/comment/p9psxf4/",
        "depth": 0,
        "score": 1,
        "author": "_Ojin",
        "excerpt": "good to see this getting measured rather than argued about. harness design ends up mattering as much as model choice and there is very little written on it.",
        "posted_at": "2026-09-14T08:45:49.000Z",
        "author_url": "https://www.reddit.com/user/_Ojin/"
      },
      {
        "url": "https://www.reddit.com/r/aiagents/comments/1wfmkvb/comment/p9qnzwm/",
        "depth": 0,
        "score": 1,
        "author": "Available_Teaching83",
        "excerpt": "This matches what I see. Same model, two harnesses, very different injection success rates, because the harness decides what the model can actually reach.\n\n The two things that move the number for me are whether tool output gets fed back as plain text into the same context as instructions, and whether the harness re-confirms a write after a tool result changes the plan mid-run.\n\n A benchmark that",
        "posted_at": "2026-09-14T12:28:05.000Z",
        "author_url": "https://www.reddit.com/user/Available_Teaching83/"
      },
      {
        "url": "https://www.reddit.com/r/aiagents/comments/1wfmkvb/comment/paxlnlp/",
        "depth": 0,
        "score": 1,
        "author": "seventyfivepupmstr",
        "excerpt": "The harness is significantly more important than the model. Not sure how people don't understand this...",
        "posted_at": "2026-09-20T10:11:56.000Z",
        "author_url": "https://www.reddit.com/user/seventyfivepupmstr/"
      }
    ],
    "evidence": [
      "[comment u/InternNo8266] This is a better way to look at security agents than raw benchmark scores. Wiz tying success rate to cost per solve makes it easier to see which setups are practical to run repeatedly"
    ],
    "link_url": "https://www.wiz.io/cyber-model-arena",
    "virality": "somewhat_high",
    "post_date": "2026-09-13T23:05:19.000Z",
    "post_kind": "link",
    "sentiment": "positive",
    "subreddit": "aiagents",
    "post_flair": [
      "News"
    ],
    "post_title": "Security benchmarks are starting to expose how much the harness matters",
    "prominence": "core",
    "source_url": "https://www.reddit.com/r/aiagents/comments/1wfmkvb/security_benchmarks_are_starting_to_expose_how/",
    "entity_role": "vendor",
    "post_author": "MushroomRight283",
    "upvote_ratio": 0.9534883720930233,
    "mention_count": 2,
    "mention_surge": false,
    "subreddit_url": "https://www.reddit.com/r/aiagents/",
    "total_upvotes": 39,
    "comments_total": 11,
    "total_comments": 11,
    "post_author_url": "https://www.reddit.com/user/MushroomRight283/",
    "signal_category": "feedback",
    "comments_included": 5,
    "products_mentioned": [
      "Cyber Model Arena"
    ]
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.