The direct answer

If you stitch together two or three upstream APIs (OpenAI, Stripe, a search provider, a payments KYC vendor), your real risk is not the headline per-minute limit. It is the combined burn rate across windows that overlap: a 60-second burst, a daily quota, and a monthly cap all resetting at different times. Forecast the budget by computing sustained requests-per-second, projecting burst headroom against your busiest minute, and alerting when actual burn exceeds a safe fraction of quota before reset. Build that, and you stop waking up to “429 Too Many Requests” pages at 3 a.m.

This matters for the founder angle directly. A quota exhaustion event is a customer-visible outage: a stuck sync, a failed checkout, a chatbot that stops replying. Recovering costs you engineering hours you do not have, plus trust you cannot rebuild easily. The cheapest fix is forecasting, not firefighting.

What “budget” actually means

A rate-limit budget is the envelope each upstream lets you consume in a defined window. Three flavors show up repeatedly:

  • Per-second or per-minute burst cap: e.g., 60 requests per minute. Short, punitive, resets fast.
  • Per-day quota: e.g., 10,000 requests per 24 hours. Resets slowly. Burns invisibly.
  • Per-month or contract ceiling: hard caps that, once hit, often require a plan upgrade or support ticket.

The standard emerging across providers (notably Cloudflare, which began shipping support in late 2025, and the IETF draft RateLimit / RateLimit-Policy headers) expresses a policy as q (quota) over w (window). A response might read RateLimit-Policy: "default";q=100;w=60 with RateLimit: "default";r=15;t=23. That is: 100 requests per 60-second window, 15 remaining, 23 seconds until the window resets.

For a small team, the practical takeaway is that each upstream defines its own policy, you can name them, and you can ask the server for the values dynamically rather than hardcoding them in your docs. The IETF draft (currently revision 11, May 2026) is the standardization track; some providers already send something close to this format.

Step 1: List every upstream with its windows

Build a single spreadsheet (or a small YAML file in your repo) with one row per upstream:

  1. Provider name and the call you make (e.g., openai.chat.completions, stripe.charges.list).
  2. Policy names you have observed in response headers, or the documented values from the vendor.
  3. Quota q and window w per policy.
  4. Cost per call: does one chat completion count as 1, or does the policy weight tokens? Some providers weight by request complexity; the IETF draft defines quota units for this reason.
  5. Reset cadence: fixed window (clock-aligned), sliding window, or token bucket.

Do not skip the cost-per-call column. Many AI and search APIs charge multiple “units” against the policy for a single call. If you ignore that, your forecast will be off by 2x to 10x and you will only learn it from a 429.

Step 2: Translate product usage into requests-per-second

Pick your busiest realistic hour, not your average. For a B2B SaaS, that might be 10 a.m. Tuesday when users log in and your backend fans out enrichment calls. For a consumer product, it might be 9 p.m. on Sunday.

For each feature path, count the upstream calls per user action:

  • User signs in: 2 calls (identity provider + profile lookup).
  • User runs the main feature: 1 LLM call, 3 search calls, 1 vector DB call.
  • Background sync every 15 minutes: 50 calls per cycle.

Multiply by expected concurrent users during your peak hour. That gives you a peak requests-per-second figure per upstream.

A worked mini-example:

  • Peak hour: 200 active users.
  • Main feature path: 1 LLM call + 3 search calls per user, spread evenly across the hour.
  • 200 users × 4 calls = 800 calls in 60 minutes = ~13.3 requests/minute averaged, with bursts at 30/minute when a batch of users hit “run” together.

Write the averaged and the burst figures next to each upstream’s quota. The gap between them is your real risk surface.

Step 3: Model burst vs. sustained

Two numbers matter:

  • Sustained rate = quota / window. This is the safe ceiling if you spread requests perfectly.
  • Burst headroom = (peak-minute requests / sustained rate per minute). Anything above 1.0x means you are bursting beyond what the policy allows even if averaged usage is fine.

For the example above, if your LLM upstream allows 60 requests per minute (sustained 1 req/sec), your 30/minute peak fits. But if 50 users hit the feature in the same minute, you are at 50 req/min burst against a 60 ceiling, leaving only 10 for background work. That is where incidents are born.

The IETF draft and Tony Finch’s analysis of it point out that fixed “quota-reset” algorithms encourage burst-then-pause behavior, which can make peaks worse than your average suggests. Linear rate-limit algorithms (e.g., GCRA) smooth the curve, and the same RateLimit headers can describe either. Read the algorithm your provider uses; it changes the math.

A simple way to model it without writing a simulator: pick your peak minute, multiply peak per-second by 60, and compare to the window quota. If the result is over ~70 percent of quota, you are in a danger zone for that single window. If it is over ~90 percent, you will hit 429s in production under realistic variance.

Step 4: Add cross-upstream correlation

A subtle failure mode: each upstream is fine on its own, but a single user action fans out across several of them simultaneously. A “generate report” button might trigger:

  • 1 LLM call (counts against token-weighted quota)
  • 4 search calls (counts against the search API)
  • 2 vector DB calls
  • 1 email-send call

When 30 users click “generate” in the same minute, no single upstream 429s in isolation, but the combined latency stretches, queues grow, and the slowest upstream becomes the bottleneck. Your rate-limit budget forecast must include this combined view, or you will tune the wrong thing.

Plot a stacked chart per peak minute: each upstream as a layer, total height as combined call volume. The layer that hits its ceiling first is your binding constraint for that minute.

Step 5: Define a burn-rate alert

Burn rate is how fast you are consuming quota relative to how fast it is replenishing. The classic formulation comes from SRE practice and translates cleanly to API quotas:

  • For a 60-second window with 100 requests, the budget refreshes at 100/60 = 1.67 requests per second.
  • If you are currently sending 3 requests per second against that policy, your burn rate is 3 / 1.67 = 1.8x.
  • A 1.0x burn rate means you will exactly hit zero at the next reset. A 2.0x burn rate means you will exhaust the window halfway through.

Practical thresholds for a small team:

  • Warning at 0.7x sustained for 5 minutes: you are trending toward exhaustion but have time.
  • Page at 1.0x sustained for 2 minutes: you will breach in this window. Throttle background jobs, defer non-critical syncs.
  • Emergency at 1.5x sustained: shed load. Return cached responses, queue user-facing work, display “slow mode” if you must.

The IETF RateLimit header gives you r (remaining) and t (seconds until reset). Compute remaining / secondsRemaining to get your current effective burn rate without instrumenting every client. Where providers do not yet ship these headers, parse whatever they do send (GitHub’s X-RateLimit-Remaining, Stripe’s 429 plus Retry-After) and compute the same number.

Step 6: Build a tiny forecast dashboard

You do not need a data warehouse. A 200-line script plus a free-tier time-series store is enough.

Minimum fields per request:

  • Timestamp
  • Upstream identifier
  • Policy name (when present in headers)
  • Remaining quota after the call
  • Reset time advertised by the server
  • Latency and status code

Compute every minute:

  • Average burn rate per upstream over the last 5 and 15 minutes.
  • Projected quota exhaustion time = remaining / current_burn_rate.
  • Days until monthly cap at current trajectory.

Display three panels: a current-state gauge per upstream, a 24-hour forecast line per upstream, and a single “binding constraint” panel showing which upstream you will hit first. That last view is what your on-call rotation actually needs.

A short checklist you can run weekly

  1. List every upstream and its documented quota windows in one file.
  2. Re-measure your peak-minute calls per upstream against that quota. Flag anything above 70 percent.
  3. Verify your client still parses whatever rate-limit headers the provider sends today.
  4. Confirm your burn-rate alert thresholds (0.7x warning, 1.0x page) actually fire on the staging load test.
  5. Check that monthly caps are not on track to be hit before month-end; if they are, slow background jobs.
  6. Review one outage postmortem from the prior week. Did the rate limit headers you relied on actually appear in the responses? If not, your parser is lying to you.

Frequently asked questions

Do I really need to forecast if my upstream has not 429ed me yet? Yes. Most quota exhaustion incidents are not first-time events; they are the moment your growth crosses a threshold you did not know was there. Forecasting finds the threshold before your customers do.

What if the provider does not send any rate-limit headers? You can still forecast by counting your outbound requests per upstream and dividing by the documented window. The forecast gets less precise (no per-request remaining count) but the budget arithmetic is the same. Treat undocumented limits as the most dangerous kind.

How is burn rate different from simple “requests per minute”? Raw rate tells you what you sent. Burn rate tells you what you sent relative to the quota being refilled. The same 30 requests/minute is fine against a 1000/minute policy and catastrophic against a 30/minute policy. Burn rate is the normalized view that lets you set one alert threshold that works across upstreams.

Should I use the new IETF RateLimit header if my provider does not send it? You cannot negotiate headers with most providers, but you can ask. Cloudflare added support in September 2025; others are watching adoption. In the meantime, build parsers that handle GitHub-style X-RateLimit-*, Stripe’s 429-plus-Retry-After, and the IETF format, and fall back to documented quotas when nothing is advertised.

What is the single highest-leverage change for a small team? Stop relying on your upstream’s 429 to be your signal. The 429 is the failure, not the warning. Pull whatever remaining-quota data you can get, compute burn rate, and page on the warning. That one change usually removes the majority of customer-visible rate-limit incidents.

A note on scope

This piece is deliberately narrow. It does not cover generic retry libraries, observability stacks, or vendor comparisons. It assumes you already know which upstreams you call and that you can read response headers. If that is true, the forecast above will tell you, in an afternoon of work, which upstream is most likely to break first and what your real burn rate is today. That is the question a small team can actually act on.

Sources