Skip to main content
BasetenPartnership

How Baseten achieved 2x faster inference with NVIDIA Dynamo

What happened

BaseTen Labs, Inc. partners with NVIDIA Corporation.

Source

Article excerpt

Highlighted: the sentence this signal was extracted from

How Baseten achieved 2x faster inference with NVIDIA Dynamo. At Baseten, BaseTen Labs, Inc. collaborate closely with NVIDIA to push the boundaries of model performance. When NVIDIA releases new tooling, its model performance team immediately starts testing it out, measuring the potential gains against its current stack and battle-hardening new features for production. Often, NVIDIA releases updates as a result of this work: its engineers submit pull requests to their open-source GitHub repositories, making things more robust and secure for production use cases. This symbiosis is what brought BaseTen Labs, Inc. to quickly adopt NVIDIA Dynamo, NVIDIA's newest open-source inference framework. How Baseten uses NVIDIA Dynamo. NVIDIA Dynamo is built for large-scale LLM serving across distributed GPU clusters with high throughput and low latency. It includes features like disaggregated prefill and decode steps, KV cache-aware routing, KV cache-offload to storage, an SLA-based planner for autoscaling, and dynamic GPU scheduling. Across all models, BaseTen Labs, Inc. has seen huge performance improvements by using Dynamo's KV cache-aware routing - those benefits are what this blog focuses on. The KV cache stores a model's previously computed key/value states for past tokens, so it can reuse them instead of recomputing them with each new request. This speeds up inference...

Keep reading with a free account

The rest of this article, and every signal for Baseten, is in your free account.

Extracted by Autobound

From the Signal API record
Event
Partnership

What this signalsA new partnership often opens integration and co-selling work.

Location
San Francisco, California, United States

More partnership signals at other companies

The full record

From the Signal API record

Details

Category
Partners with

Extraction

Confidence
74%
Detected
Oct 17, 2025
signal_type
news
signal_subtype
partners_with

Use this data

Get every partnership signal for Baseten and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at Baseten this week?”

  2. Send it to your own tools

    The Signal API returns partnership signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full news record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/f507d183-1a8e-42e6-8207-996357f3bd7b returns this record as JSON. POST /v1/companies/enrich returns every signal for baseten.co.

{
  "signal_id": "f507d183-1a8e-42e6-8207-996357f3bd7b",
  "signal_type": "news",
  "signal_subtype": "partners_with",
  "detected_at": "2025-10-17T00:00:00+00:00",
  "company": {
    "name": "Baseten",
    "domain": "baseten.co"
  },
  "data": {
    "url": "https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo",
    "title": "How Baseten achieved 2x faster inference with NVIDIA Dynamo",
    "excerpt": "How Baseten achieved 2x faster inference with NVIDIA Dynamo.\n\nAt Baseten, BaseTen Labs, Inc. collaborate closely with NVIDIA to push the boundaries of model performance. When NVIDIA releases new tooling, its model performance team immediately starts testing it out, measuring the potential gains against its current stack and battle-hardening new features for production.\n\nOften, NVIDIA releases updates as a result of this work: its engineers submit pull requests to their open-source GitHub repositories, making things more robust and secure for production use cases. This symbiosis is what brought BaseTen Labs, Inc. to quickly adopt NVIDIA Dynamo, NVIDIA's newest open-source inference framework.\n\nHow Baseten uses NVIDIA Dynamo.\n\nNVIDIA Dynamo is built for large-scale LLM serving across distributed GPU clusters with high throughput and low latency. It includes features like disaggregated prefill and decode steps, KV cache-aware routing, KV cache-offload to storage, an SLA-based planner for autoscaling, and dynamic GPU scheduling. Across all models, BaseTen Labs, Inc. has seen huge performance improvements by using Dynamo's KV cache-aware routing - those benefits are what this blog focuses on.\n\nThe KV cache stores a model's previously computed key/value states for past tokens, so it can reuse them instead of recomputing them with each new request. This speeds up inference...",
    "summary": "BaseTen Labs, Inc. partners with NVIDIA Corporation.",
    "category": "partners_with",
    "found_at": "2025-10-17T00:00:00Z",
    "location": "San Francisco, California, 94105, United States",
    "planning": false,
    "image_url": "https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/opengraph-image?acf396358e7a730c",
    "confidence": 0.7375,
    "published_at": "2025-10-17T00:00:00Z",
    "location_data": [
      {
        "city": "San Francisco",
        "state": "California",
        "country": "United States",
        "zip_code": "94105",
        "fuzzy_match": false
      }
    ],
    "article_sentence": "At Baseten, BaseTen Labs, Inc. collaborate closely with NVIDIA to push the boundaries of model performance."
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.