The decision in one paragraph

If your LLM product is early, your request volume is low, and your traffic looks more like a heartbeat than a firehose, the cheapest way to serve models is almost always a managed endpoint that bills per token and lets you scale to zero. A dedicated GPU only beats per-request managed pricing once your traffic is high enough and steady enough that the per-hour GPU rental is divided across enough useful work. Below that line, you pay for the GPU’s idle hours and you also pay cold-start pain every time the instance spins back up.

This guide walks through the math founders actually need: what “cold start” costs in latency, what “warm pool” costs in idle dollars, where the crossover sits, and how to estimate your own number without building a spreadsheet from scratch.

What cold start actually means in dollars

Cold start is the delay between a request arriving and the first token returning, when no warm container is ready. It is not one thing. It is four sequential phases that the platform has to run before your code sees a token:

  • Pulling the container image (caching helps a lot)
  • Loading model weights from storage into GPU memory
  • Capturing CUDA graphs and initializing the runtime
  • Warming the KV cache on the first requests

For a small 8B model with a warm image cache, the total cold start can land around 25 seconds. For a 70B model, the realistic floor is closer to 85 seconds on cached images, and several minutes if the image itself has to be pulled. These numbers are sourced from engineering write-ups of serverless GPU cold starts, where the phases are broken out individually rather than reported as a single vendor number.

What that means for a founder is that a “cold start is 2 seconds” claim from a marketing page is usually measuring one phase, not the whole path. If your user is staring at a spinner for 30 seconds on the first request of the day, the marketing claim did not help them.

What warm pool actually means in dollars

A warm pool is the opposite trade. You pay the platform to keep one or more instances hot, usually billed by the second of reserved capacity, so the first request after a quiet period returns in milliseconds instead of tens of seconds.

The catch is that you are now paying for idle time. If you have a single instance reserved at $2 per hour and it sits quiet for 16 hours overnight, you have spent $32 for zero user value. Multiplied across the month, an always-warm pool can quietly become your largest infrastructure line item before you have noticed.

The honest framing is: warm pool converts a latency problem into a fixed-cost bill. Cold-start managed endpoints convert a fixed-cost bill into an occasional latency spike. Each model punishes a different type of founder.

How the billing models actually behave at low volume

There are three pricing shapes you will run into:

Per-token managed (scale to zero). You pay nothing when idle. Cold starts are the platform’s problem, and they charge you only when the model is generating output. Best when traffic is bursty, sporadic, or unpredictable. The downside is that some platforms reserve the right to throttle or evict very low-volume accounts.

Per-second reserved warm pool. You pay a small hourly rate (often with a minimum commitment) to keep capacity warm. Cold starts collapse to near-zero. Best when traffic is steady enough that idle hours are mostly utilized.

Dedicated GPU rental. You pay a flat hourly rate for a specific GPU (for example, an MI300X class instance around $2 per hour in one published benchmark). You control the whole stack and can keep it warm. Best when traffic is heavy, predictable, or you need a specific model the managed providers do not host.

At very low volume, per-token managed almost always wins because the alternative is paying 24 hours of GPU time to serve 20 minutes of real traffic. As volume grows, the per-hour bill amortizes across more requests, and dedicated GPU pulls ahead once you cross a threshold that depends on average response length.

A simple crossover estimate you can run in your head

The crossover is the request rate at which the cost of a dedicated GPU equals the cost of a per-token managed endpoint. One published benchmark for a 120B-class model on an MI300X (around $2 per hour flat) versus a serverless per-token endpoint (around $0.10 per million input tokens and $0.70 per million output tokens in that test) gave the following crossovers:

  • About 18 requests per second for short responses around 30 output tokens
  • About 3 requests per second for medium responses around 220 output tokens
  • About 1 request per second for long responses around 1,200 output tokens

Those numbers are model-specific and endpoint-specific. Your mileage will vary. But the pattern is what matters: the longer each response, the lower the request rate at which dedicated GPU wins. If your product emits 1,500-token answers, you cross over surprisingly fast.

To adapt this for your own stack, you need three numbers:

  • The hourly cost of the GPU you would rent
  • The per-token price (input and output) of your managed endpoint
  • The average output tokens per request in your product

Then it is arithmetic: how many requests per hour do you need at your average output length to make the GPU’s hourly cost smaller than the token bill?

Idle timeouts and what they actually cost

“Idle timeout” sounds like an infrastructure detail. It is actually a founder decision. A managed endpoint with scale-to-zero will, after some idle window (often 5 to 15 minutes, varies by platform), tear down the warm worker. The next request pays the cold-start tax.

If your users typically arrive in clusters, an aggressive idle timeout is fine; you save money when nobody is around. If your users arrive one at a time with gaps between them, you may want either:

  • A platform or tier that lets you keep a minimum number of warm workers running
  • A pre-warm ping before the user is likely to arrive
  • A hybrid where you run on cold-start managed most of the day and switch to a warm pool during known busy windows

The dollar cost of “should I just keep one warm?” is your hourly rate times 24 hours times 30 days, minus the cold-start tax you save. At $0.10 per hour or less for a small warm worker, keeping one warm full-time is often under $75 per month. That is a sensible number for a product with paying users; it is wasteful for one with none.

When self-hosted GPU reservations beat managed per-request

Self-hosted GPU wins when three conditions are all true:

  • Traffic is predictable enough that idle time is the exception, not the rule
  • Your model, quantization, or hardware need is not served well by managed providers
  • You can absorb the operational cost of running a vLLM (or similar) service, including monitoring, restarts and capacity planning

If those three hold, a reservation or committed-use discount on a dedicated GPU usually beats per-token managed pricing once you are sustained above roughly 1 to 3 requests per second, depending on output length. Below that, you are paying for capacity you do not use.

A useful middle ground many founders miss: rent the dedicated GPU on demand, run your model warm, but burst overflow traffic to a per-token managed endpoint. You get a predictable cost floor, and the managed endpoint absorbs the spikes that would otherwise require you to over-provision.

A short decision checklist

  • Are you below a few requests per minute on average? Start with per-token managed, scale to zero. Do not pre-warm.
  • Are you between a few requests per minute and a few per second, with predictable traffic? Consider a small warm pool or a single always-on instance at the cheapest tier that fits your model.
  • Are you above a few requests per second sustained, with long outputs? A dedicated GPU or reservation probably wins. Run the crossover math with your actual average output length.
  • Do you need a model the managed providers do not host, or strict data residency? Self-host earlier than the math suggests, because the operational constraint overrides cost.
  • Are you still pre-launch with no real traffic? Per-token managed. The only bill you want at zero users is zero dollars.

FAQ

Does cold start really hurt users that much? It hurts on the first request of a session, especially for chat-style products where the user just clicked a button. For background jobs, batch processing, or anything async, cold start is usually irrelevant.

Can I just ping my endpoint every few minutes to stay warm? You can, but it is a hack. The cleanest approach is to use a platform that lets you set a minimum warm worker count and bills you only for what you reserve.

Is per-token pricing always more expensive per request than running my own GPU? For a single request, yes, you are paying a markup. The question is whether you have enough requests to justify the fixed cost of the GPU. Most early-stage products do not.

What is the single number I should track? Cost per thousand successful user requests, including the cost of cold starts you triggered and the idle time you paid for. That single ratio tells you whether your billing model still fits your traffic shape.

The founder takeaway

For an indie product in its first year, per-token managed with scale to zero is almost always the right starting point. It keeps your bill proportional to your users, not to your hopes. Move to a warm pool or dedicated GPU the day your traffic shape makes the crossover math flip, which for most LLM products happens somewhere between a handful of requests per second and a few dozen, depending on how verbose your model is.

The expensive mistake is not picking the wrong vendor. It is paying for a GPU that is idle 80% of the time while you tell yourself you are “saving money per request.”

Sources