Skip to main content
NvidiaCustomer feedback

A technical analysis by DotWave highlights that NVIDIA's reference stack for serving real-time AI models is inefficient, supporting only one concurrent session on an H100 GPU and leaving the GPU...

What happened

A technical analysis by DotWave highlights that NVIDIA's reference stack for serving real-time AI models is inefficient, supporting only one concurrent session on an H100 GPU and leaving the GPU idle approximately 75% of the time.

Source

Post

Hey guys, We've published two technical write-ups on serving real-time AI models and wanted to share the main findings here. We build the inference engine; NVIDIA develops the model we tested, Nemotron VoiceChat 11B. A request ends. A session runs on a clock. AI inference today is organized as requests: an input arrives, the model runs, the response ends and its resources are freed. If the server is busy, it can wait a moment to batch work or reorder the queue, and the cost is a slightly slower response. A growing class of models works differently. They stay active alongside something outside the GPU (a conversation, a video stream, a robot), take in input continuously and keep their state for the whole session. We call this continuous inference. Full-duplex voice is the clearest example. The model listens while it speaks, so you can interrupt it. Audio keeps arriving for the whole call, its memory of the conversation stays live until you hang up, and every 80 ms it owes the next frame of output. That deadline comes whether the server is ready or not. If the frame is late, you hear a gap. It also has to run during silence: the length of a pause is how the model tells a hesitation from the end of your turn. Unlike a classic voice agent (speech-to-text → LLM → text-to-speech, where the LLM sits idle between turns), there's no idle time to skip. Why that gets so...

Keep reading with a free account

The rest of this post, and every signal for Nvidia, is in your free account.

Extracted from these lines

  • In NVIDIA's reference stack, each 160 ms of audio becomes thousands of small GPU operations with the CPU coordinating between them, and the GPU sits idle about 75% of the time.

    From the post

  • One session fits. With two, 8–15% of audio beats arrive late. So each live conversation pays for a whole H100, most of which is waiting.

    From the post

Comments on the post

1 of 3 comments
  • “This is actually great breakdown of the batch scheduling problem. Most people don't realize how different continuous inference is from request-response pattern. I worked on similar issue last year for ASR pipeline and we kept scratching our heads why latency looked good in grafana but users complained about stutters. The 75% GPU idle time in reference stack is painful to read. That's basically b”

    u/Fabulous-Box-22902 points · Oct 3, 2026View

Extracted by Autobound

From the Signal API record
Signal
Customer feedback

What this signalsUser posts often show product pain before it reaches reviews or churn.

Subreddit
r/mlops

Companies

  • DotWaveAlso named

The full record

From the Signal API record

Numbers

Mentions
5

Details

Timing
Ongoing state
Category
Features
Virality
Very low
Post kind
Multi media
Prominence
Core
Company's role
Vendor

Topics and mentions

Topics

  • inference engine
  • gpu optimization
  • performance
  • real-time ai

Flair

  • Self-promotion :upvote:

Products named

  • Nemotron VoiceChat 11B
  • H100

Extraction

Sentiment
Negative
Detected
Oct 3, 2026
signal_type
reddit-company
signal_subtype
customerFeedback

Use this data

Get every Reddit signal for Nvidia and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at Nvidia this week?”

  2. Send it to your own tools

    The Signal API returns Reddit signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full reddit-company record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/a7762699-8199-529c-a6df-3ab169037be2 returns this record as JSON. POST /v1/companies/enrich returns every signal for nvidia.com.

{
  "signal_id": "a7762699-8199-529c-a6df-3ab169037be2",
  "signal_type": "reddit-company",
  "signal_subtype": "customerFeedback",
  "detected_at": "2026-10-03T16:07:00+00:00",
  "company": {
    "name": "Nvidia",
    "domain": "nvidia.com"
  },
  "data": {
    "nsfw": false,
    "stage": "none",
    "awards": 0,
    "timing": "ongoing_state",
    "topics": [
      "inference engine",
      "gpu optimization",
      "performance",
      "real-time ai"
    ],
    "post_id": "1wwqxbe",
    "summary": "A technical analysis by DotWave highlights that NVIDIA's reference stack for serving real-time AI models is inefficient, supporting only one concurrent session on an H100 GPU and leaving the GPU idle approximately 75% of the time.",
    "category": "features",
    "comments": [
      {
        "url": "https://www.reddit.com/r/mlops/comments/1wwqxbe/comment/pdmvct5/",
        "depth": 0,
        "score": 2,
        "author": "Fabulous-Box-2290",
        "excerpt": "This is actually great breakdown of the batch scheduling problem. Most people don't realize how different continuous inference is from request-response pattern. I worked on similar issue last year for ASR pipeline and we kept scratching our heads why latency looked good in grafana but users complained about stutters.\n\n The 75% GPU idle time in reference stack is painful to read. That's basically b",
        "posted_at": "2026-10-03T16:15:46.000Z",
        "author_url": "https://www.reddit.com/user/Fabulous-Box-2290/"
      }
    ],
    "evidence": [
      "[post] In NVIDIA's reference stack, each 160 ms of audio becomes thousands of small GPU operations with the CPU coordinating between them, and the GPU sits idle about 75% of the time.",
      "[post] One session fits. With two, 8–15% of audio beats arrive late. So each live conversation pays for a whole H100, most of which is waiting."
    ],
    "virality": "very_low",
    "post_date": "2026-10-03T16:07:00.000Z",
    "post_kind": "multi_media",
    "post_text": "Hey guys,\n\nWe've published two technical write-ups on serving real-time AI models and wanted to share the main findings here. We build the inference engine; NVIDIA develops the model we tested, Nemotron VoiceChat 11B.\n\nA request ends. A session runs on a clock.\n\nAI inference today is organized as requests: an input arrives, the model runs, the response ends and its resources are freed. If the server is busy, it can wait a moment to batch work or reorder the queue, and the cost is a slightly slower response.\n\nA growing class of models works differently. They stay active alongside something outside the GPU (a conversation, a video stream, a robot), take in input continuously and keep their state for the whole session. We call this continuous inference.\n\nFull-duplex voice is the clearest example. The model listens while it speaks, so you can interrupt it. Audio keeps arriving for the whole call, its memory of the conversation stays live until you hang up, and every 80 ms it owes the next frame of output. That deadline comes whether the server is ready or not. If the frame is late, you hear a gap.\n\nIt also has to run during silence: the length of a pause is how the model tells a hesitation from the end of your turn. Unlike a classic voice agent (speech-to-text → LLM → text-to-speech, where the LLM sits idle between turns), there's no idle time to skip.\n\nWhy that gets so...",
    "sentiment": "negative",
    "subreddit": "mlops",
    "post_flair": [
      "Self-promotion :upvote:"
    ],
    "post_title": "Why today real-time AI models cost up to 56× more than they should to serve and how fix this problem.",
    "prominence": "core",
    "source_url": "https://www.reddit.com/r/mlops/comments/1wwqxbe/why_today_realtime_ai_models_cost_up_to_56_more/",
    "entity_role": "vendor",
    "post_author": "Ok_boss_labrunz",
    "upvote_ratio": 0.5,
    "mention_count": 5,
    "mention_surge": true,
    "subreddit_url": "https://www.reddit.com/r/mlops/",
    "total_upvotes": 0,
    "comments_total": 3,
    "total_comments": 3,
    "other_companies": [
      {
        "name": "DotWave",
        "role": "alternative"
      }
    ],
    "post_author_url": "https://www.reddit.com/user/Ok_boss_labrunz/",
    "signal_category": "feedback",
    "comments_included": 1,
    "products_mentioned": [
      "Nemotron VoiceChat 11B",
      "H100"
    ]
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.