Binlog Size Calculator

July 24, 2026

MySQL and MariaDB replication log planning

Binlog Size Calculator

Estimate binary log growth from write transactions, row event size, binlog format, row image mode, changed tables, GTID overhead, DDL churn, compression, burst load, replica retention, catch-up targets, and available binlog volume.

⚙Binlog workload presets
📊Binary log sizing inputs
Committed write transactions that generate binary log events.
Approximate row payload per changed table before format and image factors.
ROW is safest for replicas and PITR, but can log more bytes.
This maps to binlog_row_image behavior for row-based logging.
Average tables touched by each transaction, including secondary write tables.
How long binlogs must remain for replicas, PITR, rebuilds, or delayed apply.
Estimated transaction payload saving from binlog transaction compression or transport compression.
Extra event headers, GTID records, checksums, metadata, and commit markers.
Schema migrations, partition rotations, online DDL, and metadata-heavy events.
Peak write multiplier for campaigns, imports, queue drains, or cache rebuilds.
Replicas or CDC consumers that must receive the binary log stream.
Target time for a replica to consume one retention window of backlog.
Usable disk reserved for active binary logs before purging or alerting.
Adds a practical shape factor for headers, checksums, and row payload patterns.
Covers measurement error, DDL variance, replica stalls, and traffic shifts.
Binlog rate
-
MB/hour
Steady binary log generation after compression and overhead.
Retention storage
-
GB retained
Binlog rate multiplied by replica retention hours.
Catch-up bandwidth
-
MB/s total
Bandwidth to drain retained backlog inside the target window.
Volume runway
-
days until full
Based on usable free space and steady binlog growth.
Enter binlog settings to calculate storage and replica risk.

Binlog calculation breakdown

Base row payload-
Format and row image factor-
Compression saving-
GTID and event overhead-
DDL contribution-
Buffered steady rate-
Peak burst rate-

Retention pressure

Volume used by retention-
Binlog files per hour at 1 GB-
Per-replica steady stream-
Total burst fan-out-
Daily binlog growth-
Recommended reserve-
🛠DB/logging spec comparison grid
MySQL 8.0GTID and ROW common

Strong replication tooling, row metadata options, transaction compression on newer releases, and crash-safe binary log indexes.

MySQL 5.7ROW default in many stacks

Often logs full row images unless changed. Check binlog_row_image, checksums, and expire settings during upgrades.

MariaDB 10.6GTID differs from MySQL

Uses MariaDB GTID semantics and replication variables; validate format compatibility before cross-family replication.

MariaDB 11.xModern retention controls

Good for current deployments, but binlog feature names and defaults can differ from Oracle MySQL documentation.

RDS MySQLManaged retention knobs

Retention may be controlled by managed procedures and backup policy. Watch storage autoscaling and replica lag together.

Aurora MySQLCluster replication layer

Binary logs are usually for external replicas or CDC. Enabling them can add measurable write and storage pressure.

Percona ServerOperational extras

Common in self-hosted estates with enhanced observability. Confirm compression and backup tooling behavior per version.

CDC sourceDebezium or Maxwell

Consumers need binlogs retained beyond connector outages. Row image minimal can break some change capture expectations.

🗃Binlog workload reference
Workload pattern Typical TPS Row event KB Binlog format Planning note
WooCommerce order bursts20 to 4003 to 12 KBROW or MIXEDOrders touch stock, sessions, metadata, coupons, and email queues, so tables per transaction matter.
Home Assistant recorder5 to 801 to 5 KBROWSmall frequent sensor writes can make GTID and commit overhead visible compared with payload size.
Small SaaS OLTP100 to 9002 to 8 KBROWReplica lag often comes from apply throughput, not raw network bandwidth.
Bulk import or migration1,000+6 to 40 KBROW or disabled on stagingShort windows can produce more retained binlog than a full quiet day.
Audit or CDC heavy updates100 to 1,5008 to 32 KBROW FULLFull row images are useful for consumers but often dominate storage.
📄Format and row image behavior
Logging mode Size tendency Replica safety Best fit Capacity caution
STATEMENTSmall when SQL is compactCan be unsafe for nondeterministic statementsLegacy simple apps and controlled batchesOne query can affect many rows while logging little, so replay behavior must be trusted.
ROW with FULL imageLargest for wide rowsStrong deterministic replicationCDC, audits, PITR, mixed app writesUpdates to wide tables or JSON columns can expand binlogs quickly.
ROW with MINIMAL imageOften much smallerGood for normal replicasHigh-write OLTP with known consumersSome CDC tools expect full before/after images or unchanged columns.
ROW with NOBLOBModerate when BLOBs are stableGood for BLOB-heavy tablesMedia metadata, CMS content, document rowsChanged BLOB or TEXT values still need to be logged.
MIXEDBetween statement and rowSafer than statement aloneCMS workloads and legacy pluginsHarder to forecast because unsafe statements switch to row logging.
📈Retention and disk runway table
Signal Low pressure Watch zone High pressure What to inspect
Retention uses volumeUnder 35%35% to 70%Over 70%binlog_expire_logs_seconds, purge policy, delayed replicas, and backup jobs.
Runway until fullOver 7 days2 to 7 daysUnder 2 daysFree volume, autoscaling, alert thresholds, and emergency purge procedure.
Daily binlog growthUnder 50 GB/day50 to 500 GB/dayOver 500 GB/dayLarge row images, DDL bursts, imports, and transaction compression.
Files per hour at 1 GBUnder 22 to 20Over 20Filesystem metadata churn, backup copy speed, and archive object count.
DDL events per dayUnder 2020 to 200Over 200Online schema changes, partition maintenance, and migration tooling.
🖥Replica catch-up and consumer table
Consumer type Uses binlog for Bandwidth driver Retention risk Planning move
Async MySQL replicaRead scale and standby recoverySteady binlog stream plus burstsReplica falls behind purge pointKeep retention beyond maintenance and network outage windows.
Delayed replicaHuman-error recovery bufferSame stream, intentionally delayed applyDelay consumes retention by designAdd delay time to outage and rebuild runway.
Cross-region replicaDisaster recoveryWAN bandwidth and packet lossBacklog grows during link issuesSize catch-up for burst drain, not only normal traffic.
CDC connectorKafka, search, analytics, cache updatesRow image width and serializationConnector outage loses source historyUse full images only when downstream systems need them.
Backup/PITR archivePoint-in-time recoveryObject copy rate and retention hoursGap breaks recovery chainMonitor archive lag and test replay to a timestamp.
💡Binlog sizing tips
Measure row image effect before changing it. binlog_row_image=MINIMAL can cut row-event volume, but some audit and CDC consumers need unchanged columns or full before/after state.
Do retention math with the slowest consumer. A single delayed replica, stopped connector, or cross-region link outage can force the source to keep binlogs far longer than the average replica needs.

Before you see replication lag, you probably see disk usage alerts. Disk usage alerts aren’t unique in that regard: they happens all the time in database management. Binary logs build up because you think they’re a low-impact log used only for recovery. They accumulate until there is too many of them on drive. Then server can no longer write because it’s out of space. Replication fails. You frantically try to revive your replicas by purging logs.

But all of this panic could of been avoided had you planned ahead by working through the storage math. Once you specify the pattern(s) by which you writes to that storage, the calculator does the math. But what’s important to know is that all of that metadata add up. Most people underestimate just how much metadata are in there. Every single transaction header gets logged. Every single row image get logged. Every single GTID record gets logged.

How to Plan for Disk Space

That’s how we do things with safety nowadays; we’re using ROW-based logging (most stacks). That tradeoff is worth it: we give up some disk space for something that we can determine will be consistent on our replicas every time. Knowing how many bytes this safety cost is key to understanding growth.

First of all, how large is the typical event? It is a few bytes for a basic log entry, which is just an updated value in a single column. It is much bigger when importing data in bulk where each row have many columns like wide JSON. Use this tool to set the average event size per row after compression and other overhead to account for true width of payload.

Secondly, what else gets touched by every transaction? You might be running WordPress/WooCommerce during some flash sale, which means every transaction touches not just stock counts but also email queues and maybe session data too. Every extra table that gets touch scales up the amount of logging. Capture that complexity instead of relying on average of a single query.

But what about the format settings? For example, if you are using CDC tools such as Debezium or Point-in-Time Recovery, then FULL row images gives you the full before-and-after state. That’s great, but it costs a lot in storage space. If your consumer doesn’t need the full picture, you can save a ton on volume size by switching to MINIMAL image mode, which logs just primary keys and changed columns. You’ll have to balance consumer compatibility vs. Storage savings here; the calculator reflects that to help you understand precisely how much tighter row image setting save you in disk space.

There’s also pressure from retention windows. Active replicas is fine with three days’ worth of logs. A crashed CDC connector? A delayed replica? Who knows if that covers them? Purging logs before any of the consumers can get up-to-date means breaking the chain of replication and losing data continuity. To avoid deleting history that some consumer will still need, the tool figures out how much overall retention storage you’ll need for the longest window where that’s possible. Abstract time becomes concrete gigabytes.

Most estimates ignore burst traffic. Yes, steady state looks fine. But databases aren’t flat; they experience spikes from cache rebuilds, report batch jobs, and campaigns. These events can double or triple there usual write throughput. You calculate your disk runway based off steady state, then suddenly a surge hits and your volume gets full sooner than expected. Burst multipliers account for these spikes in order to provide a realistic picture of worst case storage usage. It makes you size your volumes with some breathing room so you avoid those midnight alerts.

Last (but by no means least), you only have so much disk space and only so much patience to manage it as well. By calculating runways and setting good purge policies/alert thresholds, you keep things in check with a healthy system. Is it enough to survive a small outage with no human interaction? Yes. A terabyte of old log files? No. Find balance. Size the volume, plan its growth, and let it breathe. Your binary logs will remain within bounds when next flash sale occurs.

Binlog Size Calculator

Related posts

Leave a Comment