The Short Version
If you are the only person on call, monitoring and alerting are not the same job, and treating them as one is how you end up ignoring both.
Monitoring is everything your system quietly records so you can understand it later: latency, traffic, errors, resource saturation, logs, traces. Alerting is the small slice of monitoring that is allowed to interrupt you. The Google SRE book’s framing of the four golden signals (latency, traffic, errors, saturation) is useful precisely because it forces a small team to choose what matters instead of watching everything.
The practical rule for a solo founder: monitor broadly on a daily dashboard, alert narrowly on the few conditions that mean real users are hurt right now. Everything else becomes a weekly review or a “do not bother me” line in the log.
Why Mixing Them Up Burns You Out
Most early-stage teams install a monitoring tool, turn on every default alert, and then mute notifications within a week. That is not a personal failing; it is what the tooling is designed to produce. The default rules in many platforms are written for a 24/7 rotation with a runbook library, not for one person who is also answering customer emails.
A handful of sources describe the same pattern: a slow drift where dashboards get denser while the operator’s ability to act on them shrinks. The fix is not “better dashboards.” The fix is deciding, in advance, which signals deserve a page and which deserve a digest.
The Four Golden Signals, Distilled for a Small Team
Google’s SRE book defines four signals worth tracking on any user-facing system. You do not need all of them at p99 precision, but you do need to know what each one is telling you.
Latency is how long requests take to respond. The trap most small teams fall into is averaging it. Averages hide the long tail, and the long tail is what users complain about. Look at p95 or p99, and split successful and failed requests separately. A failed request that returns in 3 ms will pull your average down and make you feel great while users see error pages.
Traffic is how much demand you are taking. On its own, traffic numbers are not very useful. In context with latency and errors, they tell you whether a spike is normal load, a bot, or the start of an outage. A latency jump at 3x normal traffic means something different than the same jump at 3 AM with no traffic.
Errors are failed requests, measured as a rate rather than a raw count. Distinguish, at minimum, 4xx from 5xx. A spike in 4xx often means a client bug or a bad deploy; a spike in 5xx usually means your problem. Error budgets, which are how much failure you allow yourself before breaking your own promises, are how bigger teams keep this honest. For a one-person team, a simpler version works: decide what error rate you are willing to live with for a week, and alert when reality crosses it.
Saturation is how close you are to the ceiling on CPU, memory, disk, database connections, queue depth, or anything else with a hard limit. This is the only signal that genuinely predicts trouble before users feel it. Most teams ignore it until something else breaks first.
These four signals work because they cover the system without requiring you to instrument every line of code. You can build a credible monitoring setup from them in a weekend.
What Belongs on a Dashboard vs What Belongs in an Alert
This is the decision that decides whether you sleep.
A daily dashboard should show:
- Latency at p95 and p99, split by success and failure.
- Request rate, ideally with a comparison to the same hour last week.
- Error rate, split by class.
- Saturation on your tightest resource (often database connections or a single instance’s memory).
- A small sample of recent error logs with the noise filtered out.
An alert should fire only when:
- A user-visible SLO is at risk right now. Not “might be.” Right now.
- A resource is on a path to exhaustion within hours, not days.
- A dependency you cannot do without (auth provider, payment processor, primary database) is degraded.
- A recent deploy appears to have broken something you cannot easily roll back from.
Everything else goes into the dashboard, the weekly review, or a quiet log search. The mental model: alerting is paid for in your attention, so spend it like a founder with limited hours.
Thresholds That Survive a One-Person On-Call Schedule
Static thresholds (“alert at 80% CPU”) are easy to set and almost always wrong. They either fire during normal load and train you to ignore them, or they fire too late to matter.
A few patterns hold up better for small teams:
- Alert on user pain, not on internals. Page on error rate crossing a defined budget, not on a specific CPU number.
- Use time windows. A one-minute spike is noise; five minutes of elevated error rate is a signal. Most monitoring tools let you set “for at least 5 minutes” before paging.
- Make saturation proactive, not reactive. If your database connection pool is at 85% for an hour, you want to know on Monday morning, not when it hits 100% on Saturday.
- Separate slow-burn alerts from pager alerts. Slow-burn goes to email or a daily digest. Pager alerts go to your phone, and you should be able to count them on one hand per month.
If you cannot justify the alert with the sentence “a customer is being hurt right now and I am the only one who can fix it,” it does not belong on the pager.
The Alert-Fatigue Trap
Alert fatigue is what happens when the alerting system becomes part of the background noise you have stopped noticing. It is the single most common reason solo founders quietly stop trusting their monitoring and only find out about outages from customers.
The cure is not a smarter algorithm. It is fewer alerts. Aim for a small number of high-signal alerts, each tied to a thing you would actually do something about. When an alert fires, the right outcome should be either “fix it now,” “schedule a fix this week,” or “update the threshold because reality moved.” If none of those apply, the alert is wrong.
Tooling: How Little Is Enough
You do not need an enterprise observability platform to do this well. The categories you need to choose between are narrow:
- A uptime and synthetic check tool that watches your endpoints from outside your infrastructure.
- An application metrics and dashboard tool for the four signals.
- An error tracker that groups stack traces and tells you when a new error appears.
- An alerting layer, often built into the previous tools, that can route to email, Slack, or SMS based on severity.
For a solo founder or a small team, a single lightweight platform that covers external uptime, basic application metrics, and simple alerting is usually the right starting point. The question to ask is not “which tool has the most features” but “which tool lets me set up the four signals and a handful of alerts in under a day.” Anything heavier is borrowing complexity you do not have time to carry.
Self-hosting monitoring is possible but rarely worth the maintenance cost for a small team unless you have a hard data-residency reason. Managed services win on time-to-first-alert and on the boring work of keeping the monitoring itself up while your app is down.
A Practical Setup You Can Do This Weekend
- Pick one tool that covers external uptime checks plus basic application metrics.
- Wire the four golden signals into it: latency p95/p99 split by status, request rate, error rate by class, and saturation on the tightest resource you have.
- Put all four on a single dashboard you actually open.
- Write down, in one sentence each, the conditions under which you want a page. Keep it to three or four.
- Configure only those alerts, with a minimum 5-minute duration so you do not wake up for one-minute spikes.
- Route high-severity alerts to your phone, the rest to a daily email digest.
- After two weeks, delete every alert you did not act on.
That last step is the one most teams skip, and it is the one that makes the system survivable over months instead of weeks.
FAQ
Do I really need all four golden signals from day one?
No. Start with latency and error rate, which are the two that map most directly to user pain. Add traffic and saturation once those two are stable.
Is uptime monitoring the same as application monitoring?
Not quite. Uptime monitoring answers “is something responding?” Application monitoring answers “is what is responding actually working for users?” You usually want both, and they often live in different tools.
How many alerts should a solo founder have?
If you cannot count them on one hand, you have too many. Each alert should map to a specific action you would take.
What is an error budget in plain language?
It is the amount of failure you allow yourself over a period, expressed as a percentage. If your budget for the month is 99.9% availability, you are allowing roughly 43 minutes of downtime. Once you burn through it, the rule is usually “stop shipping features and fix reliability.”
What is the difference between monitoring and observability?
Monitoring tracks known metrics and alerts on thresholds. Observability is a broader practice of collecting enough data (metrics, logs, traces, events) that you can answer new questions you did not anticipate. For a small team, monitoring is usually enough until the system gets complex enough that you cannot predict the failure modes in advance.
Sources
- https://autoheal.ai/learn/sre-golden-signals-guide
- https://gartsolutions.com/sre-monitoring
- https://developer.cisco.com/articles/what-are-the-golden-signals/what-are-the-golden-signals-that-sre-teams-use-to-detect-issues
- https://www.splunk.com/en_us/blog/learn/sre-metrics-four-golden-signals-of-monitoring.html
- https://last9.io/blog/golden-signals-for-monitoring
- https://www.motadata.com/blog/golden-signals-monitoring-sre-metrics
- https://www.solarwinds.com/sre-best-practices/golden-signals
- https://sre.google/sre-book/monitoring-distributed-systems







