r/mlops
Notes from 22 paid runs on four GPU rental providers: what the price lists don't show
- upvotes
- 8
- comments
- 2
Post
Highlighted: the lines this signal was extracted from
I run a small service that measures what a recurring GPU job costs on rented cards, so read me as an interested party. Over three weeks we paid for 22 runs on RunPod, Vast.ai, Nebius and Hyperstack, about $37 in total. Here is what cost us time or money that no price list mentions. Startup is billed, and it varies a lot. RunPod containers were live in 5-10 seconds. The VMs on Nebius and Hyperstack took 3-7 minutes before anything ran. Nebius charged for those minutes, Hyperstack did not charge for the creating phase. Where the weights come from mattered more than I expected. Pulling 67 GB: 450 MB/s on a Vast host, 370 on Hyperstack, 210 on RunPod, 104 on Nebius. On Nebius the first video clip took 11.5 minutes, because the weights are read back from a network disk at roughly 100 MB/s. Vast containers restarted in the middle of two of our jobs, no visible cause. One job went through three restarts and lost 50 of its 166 minutes. On RunPod Community, 2 of 2 pods got evicted within 10 minutes on another job, and the minimum spot bid there was equal to the on-demand price, so we took the interruption risk without any discount. Stock: RunPod had no RTX 5090 with a CUDA 13 driver for 45 minutes, then it appeared in windows of 15 seconds to 2 minutes. Most community 5090 pods came with 46-54 GB of RAM and the job needed 64. We only caught one by polling the price API every...
Keep reading with a free account
The rest of this post, and every signal for Nebius, is in your free account.