Why client-side rate limit handling matters more than you think
If you build anything that talks to a third-party API, you’ve probably hit a 429 Too Many Requests at the worst possible time. Maybe a paying customer is staring at a loading spinner. Maybe your background job silently burned through its quota and failed twice. Maybe a chain of cron triggers all woke up at the same second and took the service out for everyone.
The good news: most rate limit pain is solved on the client side, without redesigning anything. You can build habits into your stack that detect when a limit is approaching, back off politely, and avoid stampeding a provider. This guide walks through what to look for in the response, how to retry well, and how to keep your retries from causing the next outage.
A quick caveat before we start: this article is about how your code behaves when calling someone else’s API. It does not cover designing a rate limiter for your own server. If you are the one exposing an API, that is a different problem and a different set of tools.
The two header families you will actually see
Real-world APIs do not all agree on how to tell you “slow down.” In practice, you will run into two families of headers. Treat both as part of the contract.
1. The classic Retry-After header
This is the oldest and most widely supported header. It appears on a 429 Too Many Requests response (and sometimes on 503). Its value is simple: how many seconds to wait, or an HTTP date when you can try again.
Retry-After: 30means wait 30 seconds before your next attempt.Retry-After: Wed, 21 Oct 2026 07:28:00 GMTmeans wait until that exact moment.
The Speakeasy OpenAPI documentation treats Retry-After as the standard companion to a 429 response, and the maintainer of the Ky HTTP client for JavaScript, in a Hacker News discussion of the new IETF draft, recommends that API designers “use Retry-After for as long as you can get away with it and only implementing the rate limit headers when it really becomes necessary.” That matches what you will see in production: well-run providers like Stripe, GitHub and Twilio ship Retry-After consistently.
2. The newer RateLimit and RateLimit-Policy family
The IETF has been working on a more structured header set. The current draft, draft-ietf-httpapi-ratelimit-headers-11, defines two fields:
RateLimit-Policy: a structured description of the policy in effect — its name, a partition key, the quotaq, the windowwin seconds, and the quota units.RateLimit: the policy that was actually applied to your request and what is left — the available quotarand the effective windowt.
The earlier, simpler trio — RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset — is documented at http.dev, which notes that services like GitLab, CircleCI and OKX already send it, and that it replaces the older X-RateLimit-* headers used by many APIs.
Older non-standard headers like X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset are still common in the wild. If your SDK handles one provider, it will probably need to detect several variants.
Why this matters for you
If you only react to 429s after they happen, you spend money on failed calls and frustrate users. The newer headers let you see a limit coming in a normal 200 OK response and slow down before you get blocked. That is the difference between “retry after a failure” and “shape traffic so failures rarely happen.”
Reading the headers without building a parser from scratch
You almost never need to write a structured-fields parser. The values are simple enough to handle with care.
- A bare integer is the remaining quota.
RateLimit-Remaining: 42means 42 calls left in this window. - A window expressed as
w=60means 60 seconds. - Multiple policies can appear comma-separated; treat each policy as its own budget.
- Partition keys (
pk) are hints the server gives you about which bucket you are in — per user, per app, per endpoint, or some combination. Treat them as opaque tokens; the server is supposed to document how they are generated so you can predict them.
A practical shortcut: most SDKs you would actually pay for (Ky, undici, AWS SDK, Stripe SDK, OpenAI SDK) already parse these for you. Before you build anything custom, check whether your HTTP client or SDK exposes a hook for rate limit callbacks. If it does, plug into that and save yourself a maintenance burden.
How to retry well: the four habits that prevent 90% of pain
Retry logic is where most client-side rate limit handling either shines or falls apart. Four habits cover almost every case.
1. Honor Retry-After first
If the server tells you when to retry, that is the best signal. Always prefer it over any guess you might make. This is the single highest-leverage change you can make, because the provider knows its own state better than you do.
If Retry-After is missing on a 429, fall back to a sensible default. Many APIs use 1 to 5 seconds as a polite starting point. Pick a conservative number and document it; do not pick zero.
2. Use exponential backoff for repeated failures
When you keep getting 429s after retrying — or when you hit transient 5xx errors — back off exponentially. A common pattern: 1s, 2s, 4s, 8s, capped somewhere between 30s and 60s. This gives the provider room to recover and avoids you hammering a service that is already struggling.
3. Always add jitter
This is the bit most founders forget, and it is the cause of most “every retry fires at exactly the same instant” outages. If 50 background jobs all retry after 4 seconds, you have just created a new spike.
Jitter means randomizing each retry within a window: instead of waiting exactly 4 seconds, wait somewhere between 3 and 5. Full jitter — picking a uniformly random value up to the cap — is the safest pattern and what the Ky maintainer specifically recommends for “good clients.”
4. Cap the total number of retries
Unbounded retries are how a 10-minute outage becomes a 3-hour bill. Decide in advance:
- A maximum number of attempts (often 3 to 5).
- A maximum total wall-clock time you are willing to spend.
- A clear failure mode when you give up — surface the error to the user, queue the job for later, or skip the item, depending on what the call does.
Avoiding the thundering herd
The thundering herd (or “cache stampede”) is what happens when many clients all retry at the same moment after a shared failure. It is one of the most common ways a small mistake in client code takes down a service.
The mitigation is mostly the jitter habit above, but it is worth understanding the patterns that create the problem in the first place:
- Synchronized cron jobs. If every customer’s scheduled task fires at midnight UTC, every one of them will retry at the same second. Stagger schedules by customer or by random offset.
- Workers waking from a queue pause. If your queue paused for 60 seconds and you restart 500 workers at once, they will all hit the API on the same tick. Use a startup delay with jitter.
- Cascading retries from a shared dependency. When an upstream provider recovers, every downstream caller retries simultaneously. Jitter at the client is the only realistic mitigation; coordinated “decorrelated jitter” works best.
- Reconnect storms. Clients that disconnected during an outage all reconnect together. The same fix applies.
If you operate any kind of queue, worker fleet, or scheduled-job system, build jitter and staggering in once, at the framework level. Retrofitting it later is painful.
What “good” looks like in your SDK or HTTP client
When you evaluate an HTTP client, an SDK, or an integration platform, the rate limit features that actually matter are:
- Header detection out of the box. Does it parse both
Retry-Afterand theRateLimitfamily? Does it expose the parsed values so you can log or act on them? - Built-in backoff and jitter. Or at least a clean hook so you do not have to write your own.
- Per-policy budgets. Can it track more than one policy at a time (some APIs apply different limits per endpoint or per partition)?
- Observability hooks. Does it emit events when you approach a limit, hit a 429, or back off? That is how you debug incidents without grepping logs at 2 a.m.
- Retry caps. Does it let you set a maximum attempt count and total budget?
For multi-provider integration work (one client talking to many APIs with different rules), platforms like Speakeasy, Unify, or Apideck earn their keep precisely because they centralize this logic. If you are wiring up five external APIs by hand, expect to maintain five slightly different retry strategies.
A practical starter checklist
Before you ship the next integration:
- Confirm which headers each provider actually sends. Read the docs, then check with a real
curl -i. - Make sure your client honors
Retry-Afterbefore anything else. - Add exponential backoff with jitter as the fallback path.
- Cap retries at a fixed count and a fixed wall-clock budget.
- If you run queues, schedulers, or worker pools, add staggered startup and per-job jitter.
- Log the headers you receive when you hit a 429 — they are gold for debugging provider incidents.
- Test your retry logic by simulating a 429 from a mock server. If you cannot easily test it, you do not control it.
FAQ
Do I really need to support the new RateLimit headers, or is Retry-After enough?
For most integrations today, Retry-After is enough. The Ky maintainer and several practitioners on Hacker News made the same point: ship Retry-After first, add RateLimit only when you have a real reason. The structured headers matter most when you want to shape traffic proactively rather than reactively.
What jitter strategy should I use? Full jitter (a uniformly random delay up to the cap) is the safest default. It is the pattern recommended in the IETF draft’s own discussion of retry behavior and in production SDKs. Avoid “decorrelated” or “equal” jitter unless you have measured a reason to prefer them.
How many retries are sensible? Three to five total attempts is a common range, with a hard ceiling on total wall-clock time. More than that and you are usually papering over a real problem.
Should I retry 5xx errors the same way as 429s? Often yes, with the same backoff and jitter, but never retry 4xx errors other than 429. A 401, 403, 404 or 422 will not get better by trying again.
Sources
- IETF draft: draft-ietf-httpapi-ratelimit-headers-11 (datatracker.ietf.org)
- IETF working group repository: github.com/ietf-wg-httpapi/ratelimit-headers
- Speakeasy OpenAPI documentation on rate limiting responses: speakeasy.com/openapi/responses/rate-limiting
- Hacker News discussion on HTTP RateLimit Headers: news.ycombinator.com/item?id=46618105
- http.dev expert guide to RateLimit-Limit header: http.dev/ratelimit-limit
- Tony Finch blog: HTTP RateLimit Headers (dotat.at/@/2026-01-13-http-ratelimit.html)







