SLO Compliance Calculator

September 13, 2026

HomeServerBlog reliability planner

SLO Compliance Calculator

Check whether a home lab service is inside its service level objective by combining request failures, slow responses, planned maintenance policy, downtime equivalents, error budget, and rolling burn rate.

▣SLO scenario presets

⚙SLO window and signal inputs

Used for the comparison grid and the recommended telemetry interval.
The objective for good events across this measurement window.
Use the same window as your SLO report, usually 7, 28, 30, or 90 days.
Total measured requests, probes, jobs, sync attempts, or connection checks.
HTTP 5xx, failed health checks, refused sessions, DNS failures, or job errors.
Count successful events that were too slow to be considered good.
Patch windows, router restarts, certificate swaps, or scheduled host reboots.
Unexpected downtime where the service was unavailable to normal clients.
Pick the rule you publish for your own service. Be consistent month to month.
Maximum acceptable response time for the requests counted as good.
Used as a latency margin check next to the explicit slow request count.
Short alert window used to detect fast consumption of the monthly error budget.
Total events observed in the recent burn-rate window.
Failed or too-slow events in the recent burn-rate window.
Budget intentionally held back for releases, firmware updates, and unknown faults.
Measured SLO Compliance 99.72% good events after bad-event math Compared with the selected target.
Error Budget Remaining 520 events after reserve Allowed bad events minus observed bad events.
Downtime Budget Left 93 min availability equivalent Window allowance minus counted outage minutes.
Rolling Burn Rate 0.9x recent bad-event rate vs budget rate Fast burn means the budget is draining early.

SLO formula breakdown

Budget health

Run the calculator to see compliance status.

💻Monitoring equipment/spec comparison grid

Uptime Kuma Probe30-60sHTTP, TCP, DNS, ping checks with simple alerting for home dashboards.
Blackbox Exporter15-60sPrometheus-compatible probes for HTTP status, TLS, DNS, ICMP, and TCP reachability.
Prometheus VM15sTime-series scrape engine for counters, histograms, burn-rate rules, and SLO recording rules.
Grafana Dashboard30d+Visualizes compliance, error budget, latency percentiles, and service annotations.
AlertmanagermultiRoutes slow-burn and fast-burn alerts to email, chat, webhook, or mobile push receivers.
Loki Log LabelserrorsUseful for turning status-code logs into request failure counts when app metrics are thin.
Netdata Agent1sHost-level CPU, disk, network, and container signals for explaining SLO dips.
Synthetic ClientoutsideA VPS, phone, or remote probe catches WAN, DNS, and reverse-proxy failures users feel.

For SLO math, request counters and latency histograms are better than screenshots of uptime. Use probes when the app does not expose useful metrics.

📐Derived service metrics

1200Allowed bad events

Target error rate multiplied by eligible events.

680Observed bad events

Failures plus latency misses plus counted outage equivalents.

90 msp95 latency margin

Objective minus observed p95 latency.

120Reserved events

Budget held back for releases and unexpected incidents.

📚SLO reference tables

Availability target to downtime allowance

SLO TargetError Budget30-Day DowntimeBest Home Lab Fit
99%1% bad events7 hr 12 minMedia server, dev apps, noncritical dashboards
99.5%0.5% bad events3 hr 36 minFamily file sync, photo backup, admin portals
99.9%0.1% bad events43 min 12 secVPN, reverse proxy, authentication, monitoring
99.95%0.05% bad events21 min 36 secImportant remote access with redundancy
99.99%0.01% bad events4 min 19 secDNS, critical ingress, always-on automation

Error budget burn-rate guide

Burn RateMeaningBudget ImpactTypical Action
0-1xInside planned error rateBudget lasts the full windowWatch dashboard during normal maintenance
1-2xBudget is draining slowlyMay miss target if sustainedCheck recent deploys, certificates, DNS, disks
2-6xClear reliability regressionSeveral days of budget can vanishPause changes, repair the highest-error path
6-14xFast burn incidentMonthly budget can drain quicklyPage owner, rollback, fail over, disable noisy jobs
14x+Severe user-visible outageBudget may be gone in hoursDeclare incident and prioritize restoration

Telemetry sampling and retention choices

SignalGood ResolutionRetentionSLO Use
HTTP request counter15-30 seconds35-90 daysComputes good, failed, and total event rates
Latency histogram15-30 seconds35-90 daysCounts requests beyond latency objective buckets
Synthetic probe30-60 seconds90 daysCaptures black-box availability when app metrics are missing
Incident annotationper event1 yearExplains maintenance, ISP faults, updates, and rollbacks
Host saturation1-15 seconds14-35 daysLinks SLO misses to CPU steal, RAM pressure, IO wait, or packet loss

Common home server SLO patterns

ServiceSuggested SLOLatency ObjectivePrimary Risk
Recursive DNS99.9-99.99%100 msRouter reboot, upstream resolver, cache miss storms
VPN gateway99.5-99.9%500 msISP changes, dynamic DNS, expired certificates
NAS share99-99.5%1000 msDisk scrub, ZFS resilver, sleeping disks, network saturation
Home Assistant99.5-99.9%300 msDatabase stalls, add-on restart, Zigbee or Z-Wave bridge loss
Media app98-99.5%1500 msTranscode overload, metadata scans, storage throughput

💡SLO calculation notes

Define bad events before measuring. A request-based SLO should say exactly which status codes, probe failures, timeouts, and latency misses count as bad. Changing definitions after an incident makes month-to-month reliability hard to compare.
Use burn rate for alerts, not only uptime. A service can still show a good monthly percentage while burning the whole error budget in one bad afternoon. Pair a long-window SLO chart with a short-window burn-rate alert.
This calculator treats downtime minutes as request-equivalent bad events using your average request rate. That makes request failures, latency misses, and outages comparable, but the best SLO is still the one that matches how people actually use the service.

You create a media server, spending hours tuning it. Your firewall’s tuned. Lights is automated to come on/off at the right time. Then it’s Tuesday afternoon and the router reboots. The internet drops out. Your home lab are down for 20 minutes. You get it up and running again quickly, but did you meet your reliability goals?

Most self-hosting enthusiasts don’t ask that question of themselves until it’s too late. There’s no such thing as on or off when it comes to reliability. It is something you spend every day. It’s a budget.

How to Measure Your Home Server Reliability

Here’s the tool that does all of that arithmetic for service level objectives. Remember: they’re simply reliability contracts you sign with yourself. Tell the tool what your target availability is. Tell the tool how long you measured it over (your measurement window). Then tell the tool how many times it actualy failed. The tool will tell you whether or not you’re within the lines you’ve drawn, and it’ll translate those vague percentages to real minutes of downtime.

Why does that matter? Because 99 percent uptime sounds great, until you realize that it lets you have seven hours of downtime per month. That’s totally unacceptable for something as critical as a DNS resolver; totally fine if its a personal photo archive. And the math will help you figure out where the line goes.

Keep these things in mind when inputting the data. The field for the number of eligible requests has been misinterpreted; it’s not simply all the HTTP hits. It’s also all the complete sync jobs. It’s all the successful health checks. Heck, it’s all the attempted but failed ones (which still count toward that total). If you’re only counting the successfully executed request, then your failure rate will appear artificially low.

The calculator considers not only explicit failures but also missed latencies. For example, if your server respond with a 200 OK but it takes ten seconds to do so, that’s a failure from the perspective of the user experience. Counting slow responses as your bad events provides a truer picture of the service quality, and incentivizes optimizing for speed rather than just availability.

The true magic here is error budgets. They allow you to have a target for reliability instead of perfection, because perfection is both unattainable and insanely expensive. This gives you an idea of how much room you have for error. If your target was 99.9 percent, you get some wiggle room. In fact, you burn through that wiggle room. You realize that by the first week of the month, you has already used it up, so you know you’re doing something wrong.

Now you’ve got a number from the calculator (showing you what you’ve still got) and you can see what your burn rate is (how fast you are going through that safety margin). If that’s high, then you are burning through it to quickly. It is a sign that something isn’t working correctly, and you might not notice until the entire thing is down.

Some people will trip themselves up on what counts as planned maintenance. Do I count a scheduled reboot against my SLO? Some teams don’t count any planned downtime at all, while others count it entirely because, well, the user was interrupted. You can select either option in the calculator. Either way is right (except changing the rule after-the-fact isn’t). It’s not consistent though, and consistency is needed for comparison: this month vs last month. That comparison is how your metrics help you. Without it, they’re just a bunch of numbers floating in space.

There are also more layers of complexity regarding latency goals. If your service is too slow, it doesn’t matter that it’s available. The tool will allow you to define a latency goal, and any request above that latency gets counted as failed. This is important when you’re talking about an interactive service such as a remote desktop gateway. When you’re talking about a file storage service running overnight, a couple of slow requests probably aren’t as big a deal, but it depends on the context.

Think about how your family realy uses the service. Do they watch a lot of movies? Then buffering the stream during movie night is a bad event. If the backup finishes an hour late, that might be acceptable. Maybe she’ll tolerate the backup finishing an hour later.

For a sanity check, we’ve included some reference tables that map common types of home servers to reasonable SLO targets. If your DNS resolver goes down, it breaks everything; therefore it’s worth setting a higher bar than say, a media server. Because you can defer/replace a media server (unlike DNS), it can be less reliable, it will have higher variance. Setting your target based off what actualy matters to users helps ensure you don’t spend time pursuing perfection where it won’t add value, nor ignore degradation that will cause genuine pain.

In the end, though, what compliance tracking does is force someone to be honest. It takes away the illusion that things are all good as long as the servers stay green on the dashboard. And it exposes the price we pay for instability. Now you know how your budget is being drained, so you’re able to make more informed choices regarding when to push out an update.

When do you call off an experiment that has high risk? When should you go patch these systems? What happens if we guess wrong? Can we really afford that lack of reliability?

From there, you manage reliability. Do not guess at it. Downtime should be predictable, but there will still be some downtime.

SLO Compliance Calculator

Related posts

Leave a Comment