Error Budget Burn Calculator for SLOs

September 13, 2026

HomeServerBlog reliability operations planner

Error Budget Burn Calculator

Estimate how quickly a service is consuming its SLO error budget from request failures, outage minutes, degraded minutes, rolling window length, alert threshold, telemetry source, and reserved deployment headroom.

▣Error budget burn presets

⚙SLO and observation inputs

Use combined mode when both logs and uptime probes are available.
Profile data tunes the interpretation note and comparison grid.
The reliability objective. 99.9% leaves a 0.1% error budget.
Common windows are 7, 28, 30, and 90 days.
Used for the projected-to-date context in the breakdown.
The recent window used to calculate current burn rate.
Use only requests covered by the SLI denominator.
Count user-visible 5xx, timeouts, failed probes, or failed jobs.
Minutes where the service was effectively unavailable.
Minutes with high latency, partial function, or regional impact.
A 20 minute partial incident at 25% impact counts as 5 bad minutes.
Compare current burn to your chosen multi-window alert threshold.
Hold back budget for change freezes, maintenance, and unknown incidents.
Current Burn Rate 0x observed error rate / allowed error rate Calculated from the selected SLI mode.
Budget Used 0% of total SLO budget in this lookback Lookback burn normalized to the full window.
Remaining Bad Events 0 estimated events after reserve Based on current request volume.
Time To Exhaust 0h if this burn continues Compared with your alert threshold.

Formula breakdown

Budget meter

Run the calculator to see burn status.

📡Telemetry source comparison grid

HTTP/API Logs real users Best for request availability. Denominator is total covered requests; latency depends on log shipping.
Blackbox Probe 30s-60s Best for external uptime. Denominator is probe count; catches DNS, TLS, and routing failures.
Prometheus Counters 15s pull Best for service internals. Denominator is request counter rate; align scrape gaps before paging.
RUM Beacons sessions Best for browser apps. Denominator is user sessions; sensitive to ad blockers and client sampling.
99.9% Selected SLO

Allows 0.1% bad events or bad minutes across the SLO window.

43.2m Monthly budget

For a 30 day 99.9% time-based SLO.

14.4x Fast burn

Common sharp-burn page threshold for short windows.

10% Reserve

Useful change headroom for home lab maintenance nights.

📚Error budget reference tables

SLO target versus time budget

SLO TargetAllowed Error30 Day Time BudgetUse Case
99%1.000%432 minPersonal media, non-critical lab service
99.5%0.500%216 minShared dashboard or family file service
99.9%0.100%43.2 minRemote access, DNS, password vault
99.95%0.050%21.6 minHighly visible personal cloud endpoint
99.99%0.010%4.32 minAmbitious multi-node home lab service

Burn rate alert windows

Alert TypeBurn ThresholdTypical Window PairBudget Meaning
Fast page14.4x5 min and 1 hourRoughly 2% of budget in 1 hour
Slow page6x30 min and 6 hoursRoughly 5% of budget in 6 hours
Ticket3x6 hours and 3 daysSlow leak that still needs work
Info1x1 day and 7 daysExactly spending budget pace
FreezeRemaining lowCurrent SLO windowReserve is nearly gone

Telemetry source and denominator notes

SourceGood EventBad EventPractical Limit
Reverse proxy logs2xx, 3xx, selected 4xx5xx, timeout, upstream failExcludes traffic that never reaches proxy
Service metricsHandled request counterError counter or timeout bucketNeeds stable labels and scrape health
Blackbox probeProbe successHTTP, TCP, DNS, or TLS failMay miss per-user path issues
Browser beaconSuccessful page or API spanFailed user action or long latencySampling and client blocking affect totals
Backup job logCompleted write or object putFailed job, retry exhaustion, corrupt verifyJobs are bursty, so use longer windows

Common home lab SLO examples

ServiceTypical SLOPrimary SignalWatch Closely
Home Assistant remote access99.9%HTTP probe plus API logsDNS, tunnel, certificate renewal
Jellyfin or Plex99.0%Playback start successTranscode saturation and storage stalls
Nextcloud99.5%Sync and WebDAV successDatabase locks and upload timeouts
Recursive DNS99.95%Successful query rateUpstream resolver and cache failures
VPN gateway99.9%Login and tunnel probe successISP churn and dynamic DNS lag

💡Practical burn-rate tips

Keep the denominator honest Count only user-visible traffic that your SLO promises to serve. Internal retries, health checks, and synthetic probes can hide real burn if they are mixed into the same event denominator.
Use paired windows for alerts A short window catches sharp incidents, while a longer companion window filters one-off spikes. Paging on both windows reduces noisy alerts without ignoring fast budget loss.

This calculator estimates SLO burn behavior for planning and alert tuning. Match the inputs to your own SLI definition before using the result for paging or change-freeze decisions.

Your backup script now runs every night, you’ve got 4K video streaming from your media server to your TV, and DNS resolver is pointing correctly. All lights in your dashboard is green. Then, at 10 PM on a Friday night, the whole thing turns quiet after you burned up all your reliability budget earlier in the day trying to fix some trival bug. It didn’t crash, per se, but there was no more breathing room.

In home lab ops, hobbyists treats uptime as a binary state; it’s an all-or-nothing metric where things either work or they don’t. But reliability work at large recognize uptime as a currency, a limited one at that. Each failure cost a portion of the overall amount. The calculator above allow you to monitor how quickly you’re burning through it.

Why You Need an Error Budget

An error budget is based off the notion that perfection costs money (and isn’t necessary). You decide how much downtime you want to “spend” and therefore how much downtime you can “afford.” So if you say “my service level objective is 99.9 percent,” you’re basically saying “I’m willing to accept that my service might go down for roughly forty-three minutes each month.” That’s your budget, whether you use it up for gradual degradation, accidental outages, or planned maintenance. The trick is that most hobbyist don’t do this math until they need it, but by then it’s too late, the budget is exhausted and there is panic.

#1. What will you pay for? Firstly, define exactly what failure is for you. Is a single slow-loading music stream different than your status page going down with a 500? The source doesn’t provide an example comparing a single missing file to no files being backed up.

In this calculator, there’s the option to measure time-based outages, and also just a number of failed request; which makes sense if you’re monitoring something like a web API, where request counts is important. If you’re monitoring a VPN tunnel, then uptime minutes matter more. In combined mode, the calculator take the conservative route and uses whichever metric is worse, so you have to face the worst-case scenario instead of hiding behind averages.

It also knows that time is important: a burst of errors in a short window are often more damaging than a slow leak over a month. It’s like being hit by a bucketful; it’s a sudden, broken state that requires attention now. The app does this with multiple windows of events as well, so it can separate out the emergencies from normal noise. Maybe running through alerts at 14.4x uses up a page, but maybe 3x only generates a ticket. That way you’re not waking up for every little glitch, but if something is really on fire, you wake up.

You should of also be reserving some of your budget by setting aside 10 percent or more for maintenance and deployments. This tool allow you to do that. Think of it as your emergency fund in your bank account: You hope you won’t need it, but you’re glad it’s there when you do. A single minor incident will burn through your whole month’s allowance if you haven’t reserved anything, and now you’re stuck on a freeze, no pushing any updates.

It all depends on your telemetry source. Application logs will tell you if the code are working. Blackbox probes will tell you if the server is reachable. Real user monitoring will tell you if the experience is good. So if you mix those up, you’ll have a distorted view of reality. If you want to protect the experience, use the source that matches it. The page has a clear table showing how each kind of data source aligns with different reliability goals. I think you should use that as a reference.

But ultimately, it’s all a matter of tradeoffs. Do you want something that’s fast to develop? Do you want something with high availability? You rarely get both without some price tag. But if you track your burn rate, you’ll be able to consciously decide what tradeoff to make, and no longer will you guess; you’ll know you’re safe.

When the budget gets low, you halt. And when it’s in good shape, you go. This is how to transform chaos into control, and why the stream plays on while the server remains quiet. You get to keep the weekend free.

Error Budget Burn Calculator for SLOs

Related posts

Leave a Comment