How Baseten achieved 2x faster inference with NVIDIA Dynamo
Article excerpt
Highlighted: the sentence this signal was extracted from
How Baseten achieved 2x faster inference with NVIDIA Dynamo. At Baseten, BaseTen Labs, Inc. collaborate closely with NVIDIA to push the boundaries of model performance. When NVIDIA releases new tooling, its model performance team immediately starts testing it out, measuring the potential gains against its current stack and battle-hardening new features for production. Often, NVIDIA releases updates as a result of this work: its engineers submit pull requests to their open-source GitHub repositories, making things more robust and secure for production use cases. This symbiosis is what brought BaseTen Labs, Inc. to quickly adopt NVIDIA Dynamo, NVIDIA's newest open-source inference framework. How Baseten uses NVIDIA Dynamo. NVIDIA Dynamo is built for large-scale LLM serving across distributed GPU clusters with high throughput and low latency. It includes features like disaggregated prefill and decode steps, KV cache-aware routing, KV cache-offload to storage, an SLA-based planner for autoscaling, and dynamic GPU scheduling. Across all models, BaseTen Labs, Inc. has seen huge performance improvements by using Dynamo's KV cache-aware routing - those benefits are what this blog focuses on. The KV cache stores a model's previously computed key/value states for past tokens, so it can reuse them instead of recomputing them with each new request. This speeds up inference...
Keep reading with a free account
The rest of this article, and every signal for Baseten, is in your free account.