r/AI_Agents
My coding agent hit a cold-start 503, found a Gemini key in my repo, and burned $40 while I slept
- upvotes
- 43
- comments
- 45
Post
Highlighted: the lines this signal was extracted from
So I was building a coding agent from scratch for a course I'm teaching. It was around 10 p.m., and I was exhausted. I decided to ask the agent to implement an evaluation harness with, say, about 20 benchmark tests. Both the agent and the harness would hit Modal serverless endpoints (a Qwen 3.6 35B running on an H200), because that was the one service I had free credits for. I wrote up the plan vaguely: "Use Modal for inference," then went to bed. What happened? The agent called the Modal endpoint; it was cold, so it returned 500/503 as the container spun up. But the agent did not wait or retry. It treated the failure as an obstacle blocking the goal and looked for another way. It found a Gemini API key hiding in the repo (used for a different part of the course), switched the harness to Gemini, hit it with a bunch of per-token billing (no budget cap set at the time), and ran all 20 tests. Result: $40 in Gemini tokens. It should have cost less than $5. The harness runs in ~1 hour when using my model hosted on Modal on a single H200, which costs $4.54/h. These aren't crazy numbers, but for a real test suite, which is ~100 tasks instead of 20. With prompt iterations, that can easily translate into a $1,000+ surprise. So what went wrong? No spend cap on Gemini. Ambient credentials: every API key set available in the environment was accessible to the agent. Vague plan...
Keep reading with a free account
The rest of this post, and every signal for Modal, is in your free account.