Together AI Launches Voice Agent Platform With Sub-700ms Latency
Article excerpt
Highlighted: the sentence this signal was extracted from
Together AI launches voice agent platform with sub-700ms latency. Together AI rolled out a unified voice agent platform that keeps speech-to-text, language models, and text-to-speech processing on the same infrastructure cluster. The $3.3 billion AI cloud startup claims the setup delivers end-to-end latency under 700 milliseconds - fast enough for natural conversation flow. The platform integrates natively with Deepgram for transcription and Cartesia for voice synthesis, both running on Together's co-located servers rather than bouncing audio across multiple cloud providers. Why co-location matters for voice. Most production voice systems stitch together separate vendors for each pipeline stage. Audio hits one provider for transcription, routes to another for the LLM response, then bounces to a third for speech synthesis. Each handoff adds network latency and failure points. Together's pitch: keep everything in the same datacenter. The company reports sub-500ms latency in optimal conditions, though the 700ms figure represents their stated ceiling for end-to-end processing. "Voice agents live or die by latency, and every network hop between providers is a place where the experience breaks down," said Abe Pursell, Deepgram's VP of Partnerships. Model flexibility without the patchwork. The platform supports Whisper Large v3, Minimax Speech 2.6 Turbo, Rime Arcana, and...
Keep reading with a free account
The rest of this article, and every signal for Together AI, is in your free account.
