Replication Lag Calculator

July 23, 2026

Database replication capacity planning

Replication Lag Calculator

Estimate database replica lag from primary write throughput, replica apply throughput, network latency, current lag, transaction bursts, replica fan-out, sync or async mode, catch-up target, and WAL or binlog retention.

⚙Replication workload presets
📊Lag and apply inputs
Average WAL, binlog, oplog, or redo bytes generated by the primary.
Sustained replay or apply throughput per replica after network receive.
Round-trip delay for commit acknowledgement or log shipping.
Bytes behind from pg_wal_lsn_diff, Seconds_Behind plus rate, or cloud metrics.
Optional observed replication delay; calculator also derives delay from MB backlog.
One-time spike from ETL, VACUUM-like churn, bulk load, DDL, or batch writes.
Number of replicas receiving the stream; fan-out can reduce effective headroom.
Sync modes reduce data loss exposure but add commit latency sensitivity.
Target time to clear existing lag and burst while normal writes continue.
How long logs remain available before replicas must be rebuilt or resynced.
Usable replication network throughput after TLS, VPN, compression, and overhead.
Safety factor for compression misses, indexes, conflict retries, and I/O stalls.
Lag time
-
effective current delay
Uses the larger of observed seconds and backlog-derived delay.
Catch-up ETA
-
to drain lag and burst
Only positive apply headroom can reduce lag.
Required apply
-
MB/s per slowest replica
Includes current writes, backlog, burst, overhead, and fan-out.
Retention risk
-
WAL/binlog safety
Compares lag exposure to retained log window.
Enter replication metrics to calculate lag risk.

Replication breakdown

Total backlog to replay-
Effective apply capacity-
Lag drain headroom-
Lag growth rate now-
Estimated commit latency impact-
Recommended WAL/binlog reserve-

Retention pressure

Retention used by lag exposure-
Target window status-
Network bottleneck check-
Replica fan-out note-
🛠Replication metric guide

Write MB/s

Use WAL or binlog bytes generated per second, not only SQL row size. Index updates, full-page writes, and row images can make log volume much larger than table data.

Apply MB/s

Use the slowest sustained replica apply rate. Disk IOPS, single-threaded replay, foreign key checks, long transactions, and query conflicts can reduce replay capacity.

Lag backlog

Measure bytes behind when possible. PostgreSQL LSN difference, MySQL relay log position, cloud replica lag bytes, or CDC offset size all help turn time lag into capacity math.

Retention risk

Retention is the recovery runway. If the primary discards WAL or binlogs before a replica catches up, the replica usually needs a new base backup, snapshot, or clone.

🖥Topology comparison grid

Single async replica

Simple read scale and backup target. Primary commits are fast, but replica lag can grow during bursts and failover may lose recent transactions.

Sync HA pair

Best for low data loss tolerance. Commit latency follows the replica path, so bandwidth, fsync speed, and network delay become user-facing concerns.

Multi reader fan-out

Good for read-heavy apps. Every replica must receive logs, and the slowest one often defines alert thresholds, retention reserve, and failover readiness.

Cross-region DR

Protects regional outages but adds latency and bandwidth limits. Larger WAL retention and tested rebuild plans matter more than perfect real-time apply.

📄Reference tables
PlatformLag metric to watchTypical bottleneckRetention setting
PostgreSQL streamingLSN byte diff, replay lag, flush lagWAL receive, replay I/O, conflictswal_keep_size, archive, slot retention
MySQL or MariaDBSeconds_Behind_Source, relay log positionSQL thread apply, row image volume, locksbinlog_expire_logs_seconds
MongoDB replica setOptime lag, oplog windowOplog volume, disk, secondary readsoplog size and window
SQL Server AGRedo queue, send queue, redo rateRedo thread, log send, storage latencyLog backup and truncation policy
CDC connectorsSource offset, connector lag bytesSink rate, serialization, network egressSource log retention and connector offsets
ModePrimary impactLag behaviorBest use
AsynchronousLowest commit latencyLag can grow silently if apply is slower than writesRead replicas, analytics, backups, distant DR
Semi-syncWaits for receipt by at least one replicaLess loss exposure, apply can still lagBalanced durability with moderate latency
SynchronousCommit waits for replica confirmationTime lag is small, but slow replica slows writesHA pairs and strict recovery point goals
CascadingReduces primary fan-out bandwidthDownstream replicas add another delay layerMany readers or remote branches
Logical replicationFlexible table or event stream replicationApply cost depends on transactions and indexesSelective replication, upgrades, CDC
ScenarioPrimary write rateApply targetOperational note
Small home lab database2 to 10 MB/sWrite rate plus 30% headroomRetention matters most during maintenance windows.
Busy web app OLTP40 to 150 MB/sWrite rate plus burst drainWatch long transactions and schema changes.
Analytics replica20 to 100 MB/sHigher replay during off-peakRead queries can compete with apply I/O.
Bulk import event200+ MB/s during burstSize for burst or throttle loadPre-stage imports or extend log retention.
Cross-region disaster recoveryVariableNetwork-limited plus large reserveTest failover with realistic WAN loss.
💡Replication lag tips
Alert on both bytes and seconds. Time lag alone can look small when write traffic is quiet, then become dangerous when the backlog is large and writes resume.
Size for the slowest replica. A read pool is only as safe as the replica that must survive failover. Use per-replica apply metrics instead of averages.
Protect retention during incidents. Increase WAL, binlog, or oplog retention before risky maintenance, bulk imports, network migrations, and long replica outages.
Throttle bursts when needed. If required apply rate exceeds disk or network capacity, slow the primary job, split the batch, or move heavy writes to a quieter window.

Replication lag has a way of creeping up on you. You know what I’m talking about: the replica stops getting updates, and now your reads is showing old data. At first it’s quiet. A few seconds of latency turns into minutes as the load of writes builds up. It can turn into hours if something choke the apply thread. And by the time you get an alert, there’s likely so much backlog you can’t afford to wait. That’s where replication lag feels like a ticking clock.

How much time passes before you or your users run out of patience or log history? Enter our calculator.

Why You Need a Replication Lag Calculator

Most teams always check how many seconds the secondary lags behind the primary, since this seems obvious at first but becomes harder to track once write volumes increase. In low traffic, maybe just a few KBs of unapplied writes for each second lag. But when importing large amounts of data, a ten second delay could represent gigabytes of data waiting to replay. Letting you specify how far behind in both bytes and time is what makes the difference between having a career and not. It’s a way to distinguish normal latency from structural failure.

The provided text contains no author or date information. However, there’s a second issue: the race between your retention policy (how much history gets saved) and your apply rate (how many logs can be applied per second). To conserve disk space, databases prune old binlogs or WAL files. As a consequence, when a replica lags further behind faster then the logs get deleted, it reaches a wall. There’s no catching up anymore. The primary doesn’t have history anymore! Rebuilding from scratch is generally not acceptable in a production environment. That’s where the calculator comes into play. It measures how long it would take for you to catch up compared to how long your log retention policy allows you to stay behind, which is often the point at which panic strikes.

Network overhead involves how a network deals with sending data. Does it matter how much extra weight network protocols and overhead add to your log shipping? Network protocols can create extra overhead. A cloud provider’s raw bandwidth specs don’t tell you about the replication throughput once they’re actualy trying to replicate packets. Encryption, TLS handshakes, compression misses all increase the size of each packet. To model this, the tool has a network overhead slider. Set it to ten percent or fifteen percent and this models the real-world drag that reduces speed at which you ship logs. It can make the estimated catch up time shift from twenty minutes to forty five minutes. This changes when your maintenance will finish before someone starts using the system.

The second factor is your delay tolerance, which is determined by your replication mode. Synchronous replication guarantees no data loss, but it’s slow because each transaction waits for a replica to confirm. Asynchronous replication minimizes commit latency, but it lets replicas drift out-of-sync without your knowledge. Between those two extremes, semi-sync replicates only once it has received acknowledgement from at least one replica, it asks the client to wait until then. This alters the tradeoff between consistency and availability. To account for this, the calculator adjusts its assumptions about how much latency impacts each mode. It includes the possibility that synchronous modes could slow down the primary’s performance when network latency is high.

Where many plans go wrong is with bursts. Large batch update, index rebuilds, and other ETL jobs result in sudden spikes in write volume. Steady state write rates on a replica can comfortabley handle steady state writes but stall under burst load. With the tool, you can add a one-time transaction burst. This helps you understand how much longer that extends the catch up window.

No amount of tuning helps if your needed apply rate exceeds the physical disk IOPS limits. Throttle the source or extend retention until the peak subsides.

The slowest replica is what counts. The weakest link of your read pool decides how safe your replica is. Your failover risk is as weak as the slowest node. It’s the one that gets locked by a high number of queries or by a table lock. The calculator considers that all your replicas has to catch up on the primary stream. It estimates the apply speed required for each replica to empty out the backlog in time for your target window. This allows you to know whether you should add more replicas (to spread out the load), drop some concurrent transactions (to reduce contention), or increase disk IOPS.

Your recovery runway is retention. How far back do you go? Keep it long enough to get through worst case outages. Short enough that you don’t fill up your storage with old logs you won’t use. The tool calculates and prints a retention risk score showing how much of your log window is being used at any given time due to lag. If it gets too high, you’re livig dangerously. Mess around with these ahead of time so you aren’t scrambling when it’s real.

Ultimately, “handling” replication lag is a matter of purchasing time. How much do you need? Enough to keep up with spikes but not so much that it hurts availability and/or data loss. The calculator turns vague worries into real timelines: Is today’s setup enough for tomorrow’s burst? Or should you adjust your infrastructure now while there are still logs left?

Replication Lag Calculator

Related posts

Leave a Comment