The direct answer
If you ship an API and you don’t have a number for how much unreliability your customers will tolerate, every reliability decision becomes an argument. An error budget is just that number — concrete, time-boxed, and spendable. Once you have one, you stop arguing about feelings and start looking at a graph.
This guide is for solo founders and tiny teams. It covers what an error budget is, how to calculate one from an SLO with a spreadsheet and a few metrics, and how to read burn rate without buying a six-figure observability platform.
Why this matters when you’re a team of one
When you are the only engineer, every minute counts. You don’t have an SRE team to argue with, but you do have customers who quietly leave when your API is flaky. The error budget concept gives you one specific thing: a number that tells you when to ship features and when to fix reliability.
Google’s SRE book introduced the idea: define an acceptable failure rate (the SLO), and the gap between perfect reliability and that SLO is your error budget — the failure you can ‘spend’ over a period. When the budget has room, ship freely. When it is burning fast, reliability work takes priority.
The catch: the concept is simple. The implementation varies wildly between organizations. For a solo founder, the goal is to get 80% of the value with 20% of the process. No committees, no runbooks, no JIRA tickets — just a number, a spreadsheet, and a rule you follow.
What an error budget actually is
An error budget is the allowed unreliability in your service over a defined period, usually a rolling 28 or 30 days. It is derived directly from your SLO:
- SLO = the reliability target you commit to (for example, 99.9% availability).
- Error budget = 100% minus your SLO = 0.1% in this example.
That 0.1% is the budget you can ‘spend’ on failed requests, downtime, or degraded performance before you have officially missed your SLO.
Think of it as a prepaid pain tolerance for your users. You are not promising them 100% uptime. You are promising them something slightly less, and you have a plan for what happens when you approach that line.
A concrete example
Say your API handles 1 million requests per month. With a 99.9% SLO:
- Error budget = 0.1% of 1,000,000 = 1,000 failed requests per month.
- Burn rate of 1.0 = you are consuming the budget at exactly the expected pace.
- Burn rate of 2.0 = you are failing twice as fast as allowed; you will exhaust the budget in 15 days instead of 30.
The math is trivial. The value is in what you do with the number.
How to set your first SLO (without overthinking it)
You don’t need a formal SLO document to start. You need one indicator, one target, and one time window.
Step 1: Pick one SLI
SLI = Service Level Indicator. It is the metric you measure. For most solo founders shipping an API, the first SLI is:
- Availability: (successful requests / total requests) × 100
Count 5xx errors and timeouts as failures. Count 4xx as success (the client caused it, not you). This keeps the measurement honest.
Step 2: Pick a target
Don’t start at 99.99%. That’s four nines — about 4 minutes of downtime per month. If you are running a single VPS with one database, you will miss that target constantly.
Common starting points for small teams:
- 99.9% (three nines) = ~43 minutes downtime per month. Achievable for most well-built APIs.
- 99.5% = ~3.6 hours downtime per month. Realistic for solo founders early on.
- 99% = ~7.2 hours downtime per month. Fine for internal tools or hobby projects.
Pick the lowest number you can honestly commit to without lying to customers. You can raise it later.
Step 3: Pick a window
A rolling 28 or 30 days is standard. It smooths out weekly cycles and gives you enough data to notice trends without being so long that you can’t react.
The calculation: error budget in a spreadsheet
You don’t need enterprise tooling. A Google Sheet with three columns works fine.
Column A: Total requests (rolling 30 days)
Pull this from your API logs or your hosting provider’s metrics. Most platforms — Render, Fly, Railway, Vercel — give you request counts per endpoint.
Column B: Failed requests (5xx + timeouts)
Same source, filtered for errors.
Column C: Error rate
Formula: = B / A
This is your actual reliability. Compare it to your SLO target.
Column D: Error budget remaining
Formula: = 1 - (SLO target) - C
If this goes negative, you have blown your budget for the window.
Column E: Burn rate
Formula: = C / (1 - SLO target)
A burn rate of 1.0 means you are exactly on pace. Above 1.0, you are burning faster than allowed.
Update this daily. It takes five minutes once you have the data pipeline set up.
What burn rate actually tells you
Burn rate is the speed at which you consume your error budget. It answers the question: “At this rate of failures, when will I run out of budget?”
- Burn rate < 1.0: You are under budget. You will finish the window with room to spare.
- Burn rate = 1.0: You are exactly on pace.
- Burn rate = 2.0: You will exhaust the budget in half the time (15 days instead of 30).
- Burn rate > 10.0: Something is seriously wrong. You have days, not weeks.
The Google SRE workbook recommends alerting on burn rate thresholds like 14.4× over 1 hour (which consumes 2% of the monthly budget) or 6× over 6 hours. For solo founders, a simpler rule works: if burn rate exceeds 2× for more than 24 hours, stop shipping features and investigate.
When burn rate is noise vs. a real problem
Not every spike matters. Here is how to tell the difference:
Noise (ignore it):
- A single 5xx spike that resolves in under 10 minutes and doesn’t recur.
- Burn rate above 2× for under an hour during a known deployment.
- Errors concentrated on one endpoint that you are already planning to deprecate.
Real signal (act on it):
- Burn rate above 2× sustained for 24+ hours.
- Burn rate above 5× even for a few hours.
- Errors spread across multiple endpoints with no obvious cause.
- Customer complaints arriving in your inbox or support channel.
The rule of thumb: if the graph looks bad and your customers are telling you it feels bad, believe both. If only the graph looks bad, watch it for a day before reacting.
The error budget policy: one paragraph, not a document
Enterprise teams write elaborate error budget policies. You need one sentence. Something like:
“If error budget remaining drops below 20%, I pause new features and focus on reliability until it recovers above 50%.”
Write it in your notebook. Stick it on your monitor. Follow it.
The policy exists to prevent two failure modes:
- Recklessness — shipping without measuring cumulative risk, then wondering why reliability quietly degraded.
- Over-caution — freezing all changes after a bad outage, including the fix that would have prevented it.
A solo founder who pauses feature work for a week to fix a flaky endpoint is doing the right thing. A solo founder who panics and rewrites the entire stack because of a 0.05% error spike is burning time on the wrong problem.
FAQ
Do I need an error budget if I have fewer than 100 users? Yes, but keep it simple. One SLI, one target, one spreadsheet. The value is in having a decision rule, not in the precision of the measurement.
What if my traffic is too low to calculate a meaningful rate? This is a real problem. With 100 requests per day, a single failure skews your numbers. In that case, switch from percentage-based SLOs to count-based ones (for example, “fewer than 5 failures per week”) until traffic grows.
Should I share my SLO with customers? Eventually, yes — especially if you sell to enterprises. For early-stage products, an internal SLO is enough. Public SLOs create expectations you may not want to commit to yet.
What tools do I actually need? Nothing beyond what you already have. Your hosting platform’s logs give you request counts. A spreadsheet gives you the math. A calendar reminder gives you the weekly review. That’s it.
The next step
Pick one API endpoint. Pick one SLO target (99.9% is a reasonable starting point for most). Set up a spreadsheet with the five columns above. Update it daily for two weeks. At the end of that period, you’ll know whether your reliability is actually where you think it is — and you’ll have a number that tells you when to ship and when to stop.
Sources
- Google SRE Book, Chapter 3: Embracing Risk — https://sre.google/sre-book/embracing-risk
- Google SRE Workbook, Appendix B: Example Error Budget Policy — https://sre.google/workbook/error-budget-policy
- Google SRE Workbook, Chapter 5: Alerting on SLOs — https://sre.google/workbook/alerting-on-slos
- Google Cloud Blog: How maintenance windows affect your error budget — https://cloud.google.com/blog/products/management-tools/sre-error-budgets-and-maintenance-windows







