r/mlops
Why today real-time AI models cost up to 56× more than they should to serve and how fix this problem.
- upvotes
- 0
- comments
- 3
Post
Hey guys, We've published two technical write-ups on serving real-time AI models and wanted to share the main findings here. We build the inference engine; NVIDIA develops the model we tested, Nemotron VoiceChat 11B. A request ends. A session runs on a clock. AI inference today is organized as requests: an input arrives, the model runs, the response ends and its resources are freed. If the server is busy, it can wait a moment to batch work or reorder the queue, and the cost is a slightly slower response. A growing class of models works differently. They stay active alongside something outside the GPU (a conversation, a video stream, a robot), take in input continuously and keep their state for the whole session. We call this continuous inference. Full-duplex voice is the clearest example. The model listens while it speaks, so you can interrupt it. Audio keeps arriving for the whole call, its memory of the conversation stays live until you hang up, and every 80 ms it owes the next frame of output. That deadline comes whether the server is ready or not. If the frame is late, you hear a gap. It also has to run during silence: the length of a pause is how the model tells a hesitation from the end of your turn. Unlike a classic voice agent (speech-to-text → LLM → text-to-speech, where the LLM sits idle between turns), there's no idle time to skip. Why that gets so...
Keep reading with a free account
The rest of this post, and every signal for Nvidia, is in your free account.
Extracted from these lines
In NVIDIA's reference stack, each 160 ms of audio becomes thousands of small GPU operations with the CPU coordinating between them, and the GPU sits idle about 75% of the time.
From the post
One session fits. With two, 8–15% of audio beats arrive late. So each live conversation pays for a whole H100, most of which is waiting.
From the post