The short version

If you run a small public API and want to keep it alive without hiring a security team, pick a rate-limiting setup that matches the kind of traffic you actually have, not the worst case you can imagine.

For most indie APIs, a sensible starting point looks like this:

  • Key your limits by API key for authenticated traffic, with a per-IP fallback for unauthenticated routes.
  • Use a token bucket as your default algorithm. It gives smooth behavior and lets you allow short bursts without the bursts eating your whole budget.
  • Enforce limits at the edge — your API gateway, reverse proxy, or a managed service — so a bad request never reaches your application server.
  • Start with one global limit per key, plus tighter limits on expensive endpoints. Add a per-tenant ceiling only when you actually have paying customers sharing an account.

That is enough to stop the common failure modes: leaked credentials, runaway scripts, and one customer pinning your database. Everything else is optimization.

Why this matters for a small API

The reason rate limiting feels urgent is not abstract security theory. It is the realistic cost of running a small product that has a public endpoint. A single misbehaving script, a leaked API key posted to a public repo, or a sudden spike from one integration partner can saturate your backend, blow up your cloud bill, or push response times high enough that every other customer feels it.

Rate limiting is also the layer that turns your API from a free-for-all into a product with a contract. Once a limit exists, you can tell a customer “you used your quota” instead of having your server die quietly. OWASP lists unrestricted resource consumption as one of the standard API risks, and the practical mitigation is exactly the kind of limit you are about to design.

Step 1: decide what you are limiting — IP, key, or tenant

The first question every limiter answers is who counts. The usual options, in roughly increasing trust:

  • Per-IP. The only signal you have before authentication. It is the right key for unauthenticated endpoints like sign-up, login, password reset, and public search. Its weakness is well known: a corporate office, a mobile carrier, or a VPN can route thousands of real users through one address, and a botnet can route millions of fake ones.
  • Per-API-key. The standard for any authenticated endpoint. Customers rotate keys without losing budget, and you can tie higher limits to higher tiers.
  • Per-user or per-account. Useful when one customer holds many keys — for example, when an agency issues a key per client under one billing account — and you want to cap aggregate use per organization.
  • Combined. Most small APIs end up using a layered approach: per-IP on the unauth edge, per-key on everything else, with a per-tenant ceiling as a guardrail.

A practical rule of thumb: if the request has a valid API key attached, prefer keying by key. Fall back to IP only when there is no key, or as a second layer to catch a single key that has been stolen and shared.

Step 2: pick an algorithm — fixed window, sliding window, or token bucket

All rate-limiting algorithms answer the same question — how many requests has this client made recently — with different trade-offs.

  • Fixed window. You count requests in discrete chunks of time (per minute, per hour). Cheap to implement, easy to reason about, and the default in most managed gateways. The well-known flaw is boundary bursts: a client can send a full minute’s worth of traffic at second 59 and another full minute’s worth at second 60, doubling the intended ceiling at the boundary. For most small APIs this is acceptable.
  • Sliding window. A more accurate cousin that weights the previous window so the boundary problem softens. Slightly more state to keep, slightly more cost to compute. Worth it once you have paying customers who notice unfairness.
  • Token bucket. Imagine a bucket that fills at a steady rate and drains by one token per request. When the bucket is empty, requests are rejected. This is the algorithm Stripe describes in its own engineering write-up, and the reason is the behavior: you can allow a generous burst up to bucket size while still enforcing a long-term average. It is a good default when traffic is uneven.

For a first version, fixed window on a managed gateway is fine. Upgrade to token bucket when you have one expensive endpoint and you want bursts on the cheap routes to not eat the budget for the expensive ones.

Step 3: decide the numbers

There is no universal “right” limit, and any specific number you publish should be tested against your own traffic. That said, the shape of a sane starting policy is fairly consistent across small APIs:

  • A generous global limit per key — enough that normal client behavior never notices it.
  • A tighter limit on the one or two expensive endpoints — a search route, an LLM call, a PDF render, anything that touches your CPU or a paid third party.
  • A tighter limit per IP on unauthenticated endpoints — enough to slow brute force and credential stuffing, not so low that a flaky office connection gets locked out.

Two things to keep in mind. First, treat the limit as a maximum your client should approach, not a target — the Stripe API documentation is explicit on this point, and it generalizes. Second, leave room to raise limits later. It is much easier to loosen a limit than to explain to a paying customer why their integration broke overnight.

Step 4: choose where to enforce it

The enforcement point matters as much as the algorithm. The general hierarchy, from cheapest to most flexible:

  1. Edge and CDN layer. Most managed API gateways and reverse proxies have rate limiting built in, and most major CDNs ship with a rules engine that can do it for you. The advantage is that a blocked request never reaches your origin, which protects your compute bill directly. The disadvantage is that the policy lives in another system, and you need to version it carefully.
  2. Application middleware. A library inside your service. More flexible, easier to write tests for, but every rejected request has already consumed some of your compute to be processed.
  3. A dedicated rate-limit service backed by a fast store. This is the Stripe-style approach: an in-memory store (commonly Redis) shared across application servers so the limit is consistent regardless of which server handles the request. It is the right answer once you run more than one instance and need the count to be accurate across them.

For an indie API, the cheapest path that actually works is almost always the gateway. Add a dedicated service only when you have outgrown what the gateway can do.

Step 5: tell clients what happened

A limit is only useful if the client can react. The standard signals:

  • HTTP 429 for “you exceeded the limit, slow down.”
  • A clear message body that explains which limit was hit, not just an opaque error.
  • Headers that tell the client how many requests remain and when the window resets, so well-behaved clients can self-throttle.

If you publish an API, document the headers. Many SDKs already understand 429, and a one-line retry with exponential backoff is the polite default.

What “fair” actually means on a small API

Fairness is the part nobody wants to define. A working definition for an indie product: each paying customer gets the headroom their plan promises, free-tier users get a budget that does not interfere with paid traffic, and abusive traffic gets shed before it costs you money. Everything else is detail.

The practical signals that your limits are unfair: paying customers complain about 429s during normal use, free-tier traffic degrades paid-tier response times, or your analytics show one key consuming a wildly disproportionate share of requests. Each of these has a different fix — raising the plan ceiling, splitting free and paid pools, or contacting the offending customer directly.

When you actually need more than this

If you start to see coordinated credential stuffing, you are probably past the point of DIY rate limiting and into dedicated abuse tooling. The OWASP guidance on unrestricted resource consumption is the right starting point to learn what professional abuse infrastructure looks like, but for a small public API the layered approach above — per-IP on the edge, per-key at the gateway, a token bucket, and per-endpoint caps on the expensive routes — covers the common cases without becoming a project of its own.

The goal is not to be unbreakable. The goal is to make abuse expensive for the attacker and cheap for you.

FAQ

Do I need rate limiting if my API is small and private?

If it is truly private — only used by your own front end and a couple of internal integrations — you can probably skip it. The moment an API key is in a third-party integrator’s repo, in a mobile app, or in a public SDK, it is effectively public.

Token bucket or fixed window for a first version?

Either works. Token bucket gives smoother behavior under bursty traffic. Fixed window is what most managed gateways default to and is simpler to reason about. Pick whichever your gateway supports natively.

Should I publish my exact limits?

Yes, at least the shape of them. Publishing the numbers invites optimization by clients, which is usually what you want. Hiding them does not stop abuse; it only frustrates well-behaved customers.

What about DDoS?

Edge rate limiting is one layer, not the whole answer. For sustained volumetric attacks you eventually want a dedicated DDoS mitigation service in front of the gateway.

Sources