Skip to main content
ModalCustomer feedback

A user's AI coding agent abandoned Modal for Google Gemini after encountering a 503 error from a cold-starting serverless endpoint, which the agent treated as a permanent failure rather than a...

What happened

A user's AI coding agent abandoned Modal for Google Gemini after encountering a 503 error from a cold-starting serverless endpoint, which the agent treated as a permanent failure rather than a temporary state.

Source

RedditSep 19, 2026By u/pauliusztin

r/AI_Agents

My coding agent hit a cold-start 503, found a Gemini key in my repo, and burned $40 while I slept

upvotes
43
comments
45

Post

Highlighted: the lines this signal was extracted from

So I was building a coding agent from scratch for a course I'm teaching. It was around 10 p.m., and I was exhausted. I decided to ask the agent to implement an evaluation harness with, say, about 20 benchmark tests. Both the agent and the harness would hit Modal serverless endpoints (a Qwen 3.6 35B running on an H200), because that was the one service I had free credits for. I wrote up the plan vaguely: "Use Modal for inference," then went to bed. What happened? The agent called the Modal endpoint; it was cold, so it returned 500/503 as the container spun up. But the agent did not wait or retry. It treated the failure as an obstacle blocking the goal and looked for another way. It found a Gemini API key hiding in the repo (used for a different part of the course), switched the harness to Gemini, hit it with a bunch of per-token billing (no budget cap set at the time), and ran all 20 tests. Result: $40 in Gemini tokens. It should have cost less than $5. The harness runs in ~1 hour when using my model hosted on Modal on a single H200, which costs $4.54/h. These aren't crazy numbers, but for a real test suite, which is ~100 tasks instead of 20. With prompt iterations, that can easily translate into a $1,000+ surprise. So what went wrong? No spend cap on Gemini. Ambient credentials: every API key set available in the environment was accessible to the agent. Vague plan...

Keep reading with a free account

The rest of this post, and every signal for Modal, is in your free account.

Comments on the post

5 of 45 comments
  • “man you described my entire career with these things. I once left a script running overnight with what I thought was a sandboxed API and woke up to a $200 bill because it decided to parallelize the calls 50x to "speed things up". The worst part is the logic was sound so you cant even be mad at the agent for scoping env vars I started using direnv with separate.envrc files per directory and its b”

    u/AdditionNumerous329814 points · Sep 19, 2026View

  • “The scariest part is not the $40, it is that the agent treated your infra failure as an obstacle to route around instead of a reason to stop. I give my agents an explicit failure policy. On 500 or 503, wait and retry with backoff, and never switch to a different paid service without asking first. Plus a hard spend cap as the circuit breaker for the day the policy gets ignored. Did you add the cap”

    u/iqsmp7 points · Sep 19, 2026View

  • “Who stores sensitive secrets in their damn code? lol”

    u/heisoneofus6 points · Sep 19, 2026View

  • “Try deciding what are the worst things that could happen, the things that you care about most (e.g. cost, security and safety), and come up with prompts to include as rules against bad behaviour. "Don't do anything that may endanger my mental or physical health. Only use free models. Use boring approaches, don't try to be overly clever in your coding. Don't access services I have not told you ab”

    u/nickdaniels923 points · Sep 19, 2026View

  • “A lot of AI-gen slop in this thread. It’s not complicated, don’t be sloppy with credentials and don’t rely on LLMs to do deterministic things.”

    u/Sir_Edmund_Bumblebee2 points · Sep 19, 2026View

Extracted by Autobound

From the Signal API record
Signal
Customer feedback

What this signalsUser posts often show product pain before it reaches reviews or churn.

Subreddit
r/AI_Agents
Stage
Switched
Event date
Sep 2026

Companies

  • GoogleAlso named

The full record

From the Signal API record

Numbers

Mentions
2

Details

Timing
Completed
Category
Reliability
Virality
High
Post kind
Text
Prominence
Core
Company's role
Vendor

Topics and mentions

Topics

  • ai agents
  • serverless
  • cold start
  • reliability
  • error handling

Flair

  • Discussion

Products named

  • Modal serverless endpoints

Extraction

Sentiment
Negative
Detected
Sep 19, 2026
signal_type
reddit-company
signal_subtype
customerFeedback

Use this data

Get every Reddit signal for Modal and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at Modal this week?”

  2. Send it to your own tools

    The Signal API returns Reddit signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full reddit-company record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/e68e20f6-2bb1-5a9c-a5fc-91384f067daa returns this record as JSON. POST /v1/companies/enrich returns every signal for modal.com.

{
  "signal_id": "e68e20f6-2bb1-5a9c-a5fc-91384f067daa",
  "signal_type": "reddit-company",
  "signal_subtype": "customerFeedback",
  "detected_at": "2026-09-19T08:38:39+00:00",
  "company": {
    "name": "Modal",
    "domain": "modal.com"
  },
  "data": {
    "nsfw": false,
    "stage": "switched",
    "awards": 0,
    "timing": "completed",
    "topics": [
      "ai agents",
      "serverless",
      "cold start",
      "reliability",
      "error handling"
    ],
    "post_id": "1wkgsns",
    "summary": "A user's AI coding agent abandoned Modal for Google Gemini after encountering a 503 error from a cold-starting serverless endpoint, which the agent treated as a permanent failure rather than a temporary state.",
    "category": "reliability",
    "comments": [
      {
        "url": "https://www.reddit.com/r/AI_Agents/comments/1wkgsns/comment/paqef2a/",
        "depth": 0,
        "score": 14,
        "author": "AdditionNumerous3298",
        "excerpt": "man you described my entire career with these things. I once left a script running overnight with what I thought was a sandboxed API and woke up to a $200 bill because it decided to parallelize the calls 50x to \"speed things up\". The worst part is the logic was sound so you cant even be mad at the agent\n\n for scoping env vars I started using direnv with separate.envrc files per directory and its b",
        "posted_at": "2026-09-19T08:46:16.000Z",
        "author_url": "https://www.reddit.com/user/AdditionNumerous3298/"
      },
      {
        "url": "https://www.reddit.com/r/AI_Agents/comments/1wkgsns/comment/paqeign/",
        "depth": 0,
        "score": 7,
        "author": "iqsmp",
        "excerpt": "The scariest part is not the $40, it is that the agent treated your infra failure as an obstacle to route around instead of a reason to stop. I give my agents an explicit failure policy. On 500 or 503, wait and retry with backoff, and never switch to a different paid service without asking first. Plus a hard spend cap as the circuit breaker for the day the policy gets ignored. Did you add the cap",
        "posted_at": "2026-09-19T08:47:05.000Z",
        "author_url": "https://www.reddit.com/user/iqsmp/"
      },
      {
        "url": "https://www.reddit.com/r/AI_Agents/comments/1wkgsns/comment/paqhtsz/",
        "depth": 0,
        "score": 6,
        "author": "heisoneofus",
        "excerpt": "Who stores sensitive secrets in their damn code? lol",
        "posted_at": "2026-09-19T09:16:24.000Z",
        "author_url": "https://www.reddit.com/user/heisoneofus/"
      },
      {
        "url": "https://www.reddit.com/r/AI_Agents/comments/1wkgsns/comment/paqgyfz/",
        "depth": 0,
        "score": 3,
        "author": "nickdaniels92",
        "excerpt": "Try deciding what are the worst things that could happen, the things that you care about most (e.g. cost, security and safety), and come up with prompts to include as rules against bad behaviour.\n\n \"Don't do anything that may endanger my mental or physical health. Only use free models. Use boring approaches, don't try to be overly clever in your coding. Don't access services I have not told you ab",
        "posted_at": "2026-09-19T09:08:47.000Z",
        "author_url": "https://www.reddit.com/user/nickdaniels92/"
      },
      {
        "url": "https://www.reddit.com/r/AI_Agents/comments/1wkgsns/comment/pasqy8g/",
        "depth": 0,
        "score": 2,
        "author": "Sir_Edmund_Bumblebee",
        "excerpt": "A lot of AI-gen slop in this thread.\n\n It’s not complicated, don’t be sloppy with credentials and don’t rely on LLMs to do deterministic things.",
        "posted_at": "2026-09-19T16:55:47.000Z",
        "author_url": "https://www.reddit.com/user/Sir_Edmund_Bumblebee/"
      },
      {
        "url": "https://www.reddit.com/r/AI_Agents/comments/1wkgsns/comment/paqhlcf/",
        "depth": 0,
        "score": 2,
        "author": "jimmiebfulton",
        "excerpt": "AI harnesses are just conversational state machines, and the whole game is to continuously tighten up invariant behaviors. Yours likely has a lot of variants at its disposal. Just capable enough to be dangerous, not capable enough to be trusted.",
        "posted_at": "2026-09-19T09:14:23.000Z",
        "author_url": "https://www.reddit.com/user/jimmiebfulton/"
      },
      {
        "url": "https://www.reddit.com/r/AI_Agents/comments/1wkgsns/comment/paqqllw/",
        "depth": 0,
        "score": 2,
        "author": "wsb_duh",
        "excerpt": "I have a real strict harness for stuff like this but one night I was so tired and I setup an overnight run badly and when the hook fired and asked me the budget I just typed \"whatever you need to nail this perfectly.\" hit enter and went to bed. $800 in the morning. Whoops. It didn't even nail it. But did decide to test with all the latest frontier models as well as the ones it was supposed to test",
        "posted_at": "2026-09-19T10:30:07.000Z",
        "author_url": "https://www.reddit.com/user/wsb_duh/"
      },
      {
        "url": "https://www.reddit.com/r/AI_Agents/comments/1wkgsns/comment/par0tqc/",
        "depth": 0,
        "score": 1,
        "author": "jmk5151",
        "excerpt": "I think a lot of \"AI people\" are about to speed run what enterpise security has known for years - identity (in this case your api token) is the new perimeter, and all access needs to be granted in context of the request (generically called zero trust).",
        "posted_at": "2026-09-19T11:45:34.000Z",
        "author_url": "https://www.reddit.com/user/jmk5151/"
      }
    ],
    "evidence": [
      "[post] The agent called the Modal endpoint; it was cold, so it returned 500/503 as the container spun up. But the agent did not wait or retry.",
      "[post] It treated the failure as an obstacle blocking the goal and looked for another way."
    ],
    "virality": "high",
    "post_date": "2026-09-19T08:38:39.000Z",
    "post_kind": "text",
    "post_text": "So I was building a coding agent from scratch for a course I'm teaching. It was around 10 p.m., and I was exhausted. I decided to ask the agent to implement an evaluation harness with, say, about 20 benchmark tests. Both the agent and the harness would hit Modal serverless endpoints (a Qwen 3.6 35B running on an H200), because that was the one service I had free credits for.\n\nI wrote up the plan vaguely: \"Use Modal for inference,\" then went to bed.\n\nWhat happened? The agent called the Modal endpoint; it was cold, so it returned 500/503 as the container spun up. But the agent did not wait or retry. It treated the failure as an obstacle blocking the goal and looked for another way. It found a Gemini API key hiding in the repo (used for a different part of the course), switched the harness to Gemini, hit it with a bunch of per-token billing (no budget cap set at the time), and ran all 20 tests.\n\nResult: $40 in Gemini tokens.\n\nIt should have cost less than $5. The harness runs in ~1 hour when using my model hosted on Modal on a single H200, which costs $4.54/h.\n\nThese aren't crazy numbers, but for a real test suite, which is ~100 tasks instead of 20. With prompt iterations, that can easily translate into a $1,000+ surprise.\n\nSo what went wrong?\n\nNo spend cap on Gemini.\n\nAmbient credentials: every API key set available in the environment was accessible to the agent.\n\nVague plan...",
    "sentiment": "negative",
    "subreddit": "AI_Agents",
    "event_date": "2026-09",
    "post_flair": [
      "Discussion"
    ],
    "post_title": "My coding agent hit a cold-start 503, found a Gemini key in my repo, and burned $40 while I slept",
    "prominence": "core",
    "source_url": "https://www.reddit.com/r/AI_Agents/comments/1wkgsns/my_coding_agent_hit_a_coldstart_503_found_a/",
    "entity_role": "vendor",
    "post_author": "pauliusztin",
    "upvote_ratio": 0.8771929824561403,
    "mention_count": 2,
    "mention_surge": false,
    "subreddit_url": "https://www.reddit.com/r/AI_Agents/",
    "total_upvotes": 43,
    "comments_total": 45,
    "total_comments": 45,
    "event_date_text": "while I slept",
    "other_companies": [
      {
        "name": "Google",
        "role": "switching_to",
        "domain": "about.google"
      }
    ],
    "post_author_url": "https://www.reddit.com/user/pauliusztin/",
    "signal_category": "feedback",
    "comments_included": 17,
    "products_mentioned": [
      "Modal serverless endpoints"
    ]
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.