The short version

When your service goes down and you are the entire team, the goal of the first 30 minutes is not to fix the bug. The goal is to stop bleeding trust. Pin the incident, tell the right people in the right order, and start a single document you will keep updating. Fixes come after. The framework below is built around three time-boxed phases — triage (0–10 min), stabilize (10–30 min), and recover and learn (30 min and beyond) — followed by a blameless postmortem template you can actually fill in while the memory is fresh.

If you only have five minutes, scroll to the 30-minute timeline and the postmortem template. Everything else is context for why those steps exist.


Why a solo founder needs a runbook at all

Most runbook templates assume there is someone else to page. A solo founder, or a tiny team of two or three, does not have on-call rotations, a communications lead, or a separate incident commander. The whole role collapses onto one person, often in the middle of the night, often while customers are already tweeting.

That is exactly when a runbook earns its keep. A runbook is not a novel. It is a short, opinionated checklist that turns panic into a sequence of small decisions. The single biggest variable in your mean time to recovery is not your stack — it is whether the next action is obvious or improvised.

The tradeoff is real: writing a runbook takes a calm afternoon you do not feel you have. The honest answer is that you will trade one calm afternoon for several future evenings that are not calm. The runbook below is designed to be written in a single sitting and reused forever.


What you need before the first incident

You do not need enterprise tooling. You need four things in place before anything breaks:

  1. A single incident document. A plain text file, a Notion page, or a Google Doc with a name like inc-2025-01-14-checkout-500s. The medium matters less than the rule that there is exactly one of them and it has a timestamped timeline.
  2. A status page. Even a free tier from providers in this category gives customers somewhere to look besides your inbox. If you cannot set one up today, a pinned post on your social account is a workable substitute — but a real status page is meaningfully better because it does not require your active attention to update.
  3. A short comms tree. Three to five email addresses or chat handles: you, a backup, a co-founder if you have one, a vendor contact for your hosting or payments provider, and one person who is allowed to publicly answer customers when you are head-down debugging.
  4. A list of golden signals for each service. For every service you run, write down what “healthy” looks like: error rate, latency, queue depth, disk, and one synthetic check you can hit from your phone. You will use these in the triage phase.

The goal of these four artifacts is to remove thinking during the incident. If you find yourself asking “what is the URL of my dashboard again?” at 2 a.m., the runbook has already failed.


The 30-minute timeline

Below is the exact sequence. Times are ceilings, not targets — if you stabilize faster, skip ahead.

Minutes 0–10: Triage and declare

  • Acknowledge. Open the incident document the moment you suspect an incident, even if you are not sure. A suspected incident that turns out to be a false alarm is cheap. An undeclared outage is expensive.
  • Assess severity. A simple four-level scale works:
    • SEV1: complete outage or data loss. Drop everything.
    • SEV2: major feature broken for most users.
    • SEV3: minor degradation, workaround exists.
    • SEV4: cosmetic, internal-only, or single-user.
  • Answer four questions: What is impacted? How many users? Is impact still spreading? Is revenue or compliance affected? These answers determine SEV level and how often you update people next.
  • Notify your inner circle. Send a one-line message to your comms tree: incident declared, SEV level, current best guess at impact. No detail, no jargon, no apologies — just the headline.
  • Open the comms channel. A single chat thread, ideally named after the incident, where every decision and observation is logged. If you are solo, this can be a notebook. If you have a team, it must be a shared channel that is not your personal DMs.

Minutes 10–30: Stabilize

  • Stop the bleeding before you find the cause. Roll back the last deploy if the incident started within an hour of a release. Failover the dependency. Disable the misbehaving feature flag. Throttle the abusive endpoint. None of these require root cause.
  • Check the golden signals. Look at error rate, latency, saturation, and traffic. One of them will be obviously wrong, and that is your first hypothesis, not your conclusion.
  • Update the status page. Even a short message — “We are investigating elevated errors on checkout” — is better than silence. Customers can wait; silence makes them angry.
  • Set the update cadence. SEV1 gets an update every 10–15 minutes, even if the update is “still investigating.” SEV2 every 30 minutes. SEV3 whenever something changes.
  • Keep a timeline. Every action, every observation, every external message — timestamped. You will need this for the postmortem and possibly for customer support tickets later.

After 30 minutes: Recover, watch, and write it up

  • Confirm the fix. Check error rate is back to baseline, latency is normal, and synthetic checks pass. “I think it’s fixed” is not confirmation. Numbers are.
  • Watch for regression for 30–60 minutes. Most incidents recur within the first hour after a fix because the underlying condition has not been removed, only masked.
  • Post a resolution note. “Resolved at HH:MM. Root cause: X. We are writing a full postmortem and will share it by [date].”
  • Close the incident. Mark the document as closed, but do not archive it. It becomes input to the postmortem.

How to communicate when you are the only voice

The hardest part of a solo incident is not the debugging — it is deciding what to say and how often. Three rules carry most of the weight.

Rule 1: Tell people what you know, what you do not know, and what you are doing. “Checkout is returning 500 errors for an unknown number of users. We are investigating. Next update in 15 minutes.” That sentence works for SEV1 and SEV2 alike. It is honest, short, and creates a clear expectation.

Rule 2: Never promise a time you are not certain of. “Fixed in 30 minutes” becomes a lie the moment it slips. “We are working on it and will update every 15 minutes” is a promise you can keep even on a bad night.

Rule 3: One channel, one voice. Pick the status page or your support inbox — not both, not all of them. Customers should know where to look. If you delegate public replies to someone else, brief them in plain language and tell them to say “we” not “I.”


Status page setup for a small business

A status page is one of the cheapest reliability wins available. The category includes hosted services with free tiers and self-hosted options. Either path works; the difference is mostly about how much you want to think about it.

When evaluating a status page tool, look for:

  • Free or near-free tier. You should not be paying real money until you have paying customers who care about SLAs.
  • Incident templates. Pre-written messages for “investigating,” “identified,” “monitoring,” “resolved” save you from composing under stress.
  • Email and webhook subscriptions. Customers should be able to subscribe to updates without logging in.
  • Custom domain support. status.yourcompany.com looks more credible than a third-party subdomain.
  • Component-level status. If you have more than one service, you want each one on its own line.

A setup decision worth thinking through up front: whether to subscribe customers automatically when they sign up for your product, or to leave it opt-in. Auto-subscribe reduces inbound “is it down?” tickets during incidents. Opt-in is friendlier to privacy-conscious users. Most small products land on opt-in with a prominent link in the footer and a one-line mention in the welcome email.


The blameless postmortem template that actually prevents the next incident

A postmortem is only useful if it changes behavior. Most postmortems fail because they are either too vague (“we need better monitoring”) or too personal (“Alice missed the alert”). The template below forces specificity without blame.

Section 1: Summary

  • Incident ID and title.
  • Date and time of detection, time of resolution, total duration.
  • SEV level (final, not initial).
  • One-paragraph customer-facing summary.

Section 2: Timeline

A bulleted list of every meaningful event with timestamps. Pull this directly from the incident document. Format: HH:MM — [action or observation] — [who, if relevant].

Section 3: Impact

  • Number of users affected (estimate is fine; “all” is fine).
  • Revenue impact, if any.
  • Support tickets generated.
  • Any data loss or compliance exposure.

Section 4: Root cause

Write this in plain language a non-engineer could follow. “A deploy at 14:32 introduced a database connection leak. Under load, connections were exhausted within 12 minutes, causing 500 errors on all write paths.” No jargon for jargon’s sake.

Distinguish carefully between proximate cause (the immediate trigger), contributing factors (the conditions that let it happen), and root cause (the thing that, if fixed, would have prevented it). Most incidents have one of each.

Section 5: What went well

This section is not optional. List three to five things that worked — a dashboard that caught it early, a rollback that held, a teammate who asked the right question. Without it, the postmortem becomes a list of failures and people stop reading.

Section 6: What went poorly

Likewise, three to five things. Phrased as systems or processes, not people. “The alert fired but did not page anyone” is a system observation. “Nobody noticed the alert” is a people observation and tends to produce defensiveness.

Section 7: Action items

Each action item needs four fields: what will change, who owns it, by when, and how we will know it is done. An action item without an owner is a wish. An action item without a deadline is a wish with a date.

Limit yourself to three to five action items. More than that and nothing gets done. If you have more than five, rank them and defer the rest to a “next time” list.

Section 8: Lessons

One paragraph on what this incident taught you about your service, your tooling, or your response process. This is the section you re-read six months later.


How to make the postmortem actually blameless

“Blameless” does not mean “no accountability.” It means the postmortem describes what happened and why, not who is at fault. In practice:

  • Refer to people by role, not name, when describing actions (“the on-call,” “the deployer”). Names go in the action-item ownership field only.
  • Assume the person who was involved made the best decision they could with the information they had. If they did not, the question is what information was missing, not what they were thinking.
  • Never share the postmortem publicly with names attached. Internally, names are fine. Externally, role-only.

The reason this matters more for solo founders than for big teams is that the “person” is you. A blameless postmortem is how you forgive yourself for a mistake and still learn from it.


Where to keep the runbook and postmortems

Your runbook should live somewhere searchable, version-controlled, and accessible from your phone. A markdown file in a private repo, a Notion workspace, or a dedicated docs tool all work. What does not work: a Slack pinned message, a sticky note, or your memory.

Your postmortems should live in the same place, tagged by date or service. The goal is that when the next similar incident happens, you can search past postmortems in under a minute. If you cannot do that, the postmortems are not earning their cost.


A short FAQ

How long should an incident response runbook be? Long enough to cover the four phases above, short enough to read in 10 minutes during an incident. For a solo founder, that is usually one to two pages. Anything longer is reference material, not a runbook.

Do I really need a status page if I am pre-revenue? You do not need a fancy one. A free tier from a hosted status page provider, or even a public Notion page that you update by hand, is enough. What you do need is a single, canonical place where the answer to “is it down?” is not your inbox.

What is the difference between a runbook and a postmortem? A runbook is used during the incident to guide action. A postmortem is written after the incident to extract learning. They live in the same doc system but serve opposite purposes.

How often should I update the runbook? After every incident that revealed a gap. A reasonable cadence for a solo founder is once per quarter, plus immediately after any incident where the runbook would not have helped.

Should I share postmortems publicly? Often, yes. Public postmortems build trust with customers and force a quality bar on the writing. Strip names and any customer-specific detail first.


Sources