Write Amplification Calculator

July 23, 2026

Database storage and SSD endurance planning

Write Amplification Calculator

Estimate physical SSD writes from logical database writes, storage engine behavior, index fanout, compaction, replication, WAL or binlog overhead, page rewrites, erase block overhead, and compression.

⚙DB and storage presets
📊Write amplification inputs
Application-level changed data before database logs, indexes, compaction, and SSD effects.
Preset engine behavior sets typical base write overhead and compaction posture.
Indexes touched by each write. More indexes usually mean more pages and more WAL.
Extra rewritten bytes from LSM compaction, merge trees, vacuum rewrite, or COW churn.
Storage-side copies that also land on SSDs. Use 1 for a single primary device.
Additional log bytes relative to logical writes, including WAL, redo, binlog, AOF, or journal.
Extra page or extent rewrites from random updates, copy-on-write, full pages, and fragmentation.
Flash translation layer garbage collection, write coalescing, spare area, and erase block mismatch.
Physical stored size divided by logical size. Use 1.00 for no compression, 0.50 for 2:1.
Drive or array write endurance rating. Results show estimated percent consumed per day.
Your desired amplification ceiling for optimization planning.
Lower reserve raises FTL and compaction pressure. Keep SSDs and LSM stores with breathing room.
Physical writes
-
TB per day
Total estimated bytes written to SSD media.
Amplification factor
-
physical / logical
Combined engine, log, index, COW, and flash overhead.
SSD wear
-
endurance per day
Daily TBW budget consumed by this workload.
Optimization target
-
GB/day reduction
Write reduction needed to meet the target WAF.

Amplification breakdown

Compressed logical base-
Index write overhead-
WAL / binlog writes-
Compaction / merge writes-
Page rewrite writes-
Erase block overhead-
Replicated physical writes-

Wear posture

Run the calculator to see the write endurance posture.
Estimated monthly writes-
Estimated yearly writes-
Endurance runway-
Biggest optimization lever-
🛠Current engine factor snapshot
Heap
Engine family

B-tree indexes and heap page rewrites dominate many OLTP systems.

1.10x
Base overhead

Internal metadata, record headers, checksums, and allocator behavior before user knobs.

0.18x
Per index

Extra write factor per secondary index touched by a typical write path.

3-8x
Common WAF

Field ranges vary by workload shape, update locality, free space, and durability settings.

🗃Engine comparison grid
Engine or storage layer Write pattern Common WAF range Main amplifier Best first lever
PostgreSQL heap + B-treeWAL plus heap and indexes3x to 8xIndexes, full pages, vacuum churnHOT updates, index cleanup, checkpoint tuning
MySQL InnoDBRedo, doublewrite, clustered pages3x to 7xSecondary indexes and page splitsIndex pruning, fill factor, redo sizing
RocksDB / LevelDBAppend memtables, flush, compact5x to 20xLevel compaction and tombstonesCompaction style, level sizes, compression
Cassandra / ScyllaDBCommitlog plus SSTable compaction4x to 16xCompaction and repair churnCompaction strategy and TTL hygiene
MongoDB WiredTigerJournal plus B-tree pages3x to 10xDocument growth and index fanoutSchema shape, index set, compression
ClickHouse MergeTreePart writes then background merges2x to 12xSmall parts and merge backlogBatch size, part count, merge settings
Kafka append logSequential append and replication1.2x to 4xReplication and segment cleaningRetention, compression, replication factor
ZFS copy-on-writeCOW blocks plus intent log2x to 8xRecordsize mismatch and snapshotsRecordsize, sync policy, snapshot pruning
📈Write amplification formula table
Component Calculator term What it represents When it rises Typical mitigation
Compressed baseLogical writes x compressionUser data after compression or expansionPoorly compressible payloads, encrypted blobsColumnar formats, payload trimming, better codecs
Index fanoutIndexes x per-index factorSecondary index pages and metadataMany indexes, random keys, wide indexed valuesDrop unused indexes, use narrower keys
WAL or binlogBase x WAL factorDurability log, redo, journal, binlog, AOFFull page writes, row images, fsync batchingBatch writes, tune log compression safely
CompactionBase x compaction factorRewrites from LSM levels, merge trees, vacuumTombstones, tiny SSTables, merge backlogBatch ingest, tune levels, clear tombstones
Page rewriteBase x rewrite percentCopy-on-write blocks, page splits, extent churnRandom updates, low fill space, snapshotsFillfactor, recordsize, update locality
SSD erase blocksSubtotal x erase overheadFTL garbage collection and erase block mismatchLow free space, mixed random writes, old SSDsOverprovisioning, TRIM, free space reserve
ReplicationSubtotal x replication factorCopies written across local mirrored or replicated mediaRF 3 clusters, mirrored writes, sync replicasSeparate endurance planning per replica tier
📝Common DB and storage presets
Preset Logical writes Indexes Compaction WAL / log Compression
PostgreSQL OLTP500 GB/day40.4x0.70x0.65
MySQL InnoDB SaaS650 GB/day50.3x0.85x0.70
RocksDB LSM Store900 GB/day24.5x0.45x0.45
Cassandra Wide Rows1200 GB/day13.2x0.35x0.55
MongoDB WiredTiger700 GB/day60.8x0.60x0.50
SQLite Edge Node40 GB/day30.2x0.90x0.95
ClickHouse MergeTree2500 GB/day12.4x0.12x0.32
Kafka Log Broker1800 GB/day00.2x0.15x0.45
Ceph BlueStore OSD900 GB/day11.1x0.40x0.75
ZFS Home NAS250 GB/day00.6x0.25x0.60
Redis AOF Heavy300 GB/day00.4x1.20x0.80
Prometheus TSDB350 GB/day11.6x0.25x0.38
🔧SSD wear reference table
Daily wear rate Endurance posture Runway signal Operational response What to measure
Under 0.05% / dayLowMore than 5 yearsKeep monitoring and preserve free spaceSMART host writes and media writes
0.05% to 0.15% / dayModerateAbout 2 to 5 yearsReview index count, compaction, and write burstsDatabase WAL, compaction metrics, iostat
0.15% to 0.35% / dayHighAbout 10 to 24 monthsPrioritize write reduction or higher endurance mediaFTL writes, erase counts, write latency
Over 0.35% / dayCriticalUnder 10 monthsAct quickly: lower WAF, shard writes, or replace media classWear leveling count, spare blocks, throttling
💡Write amplification optimization tips
Start with the biggest multiplier A single unused secondary index can cost more than a small WAL setting change on update-heavy tables.
Keep free space available SSD garbage collection, LSM compaction, and copy-on-write filesystems all behave better with headroom.
Batch small writes carefully Larger batches reduce per-transaction log overhead, but very large batches can cause merge and checkpoint spikes.
Compare database and device counters Use WAL, compaction, and engine metrics alongside SMART host writes and media writes to estimate real WAF.

If you purchased an enterprise SSD because of its high capacity and blazingly fast reads, there’s a good chance you’re overlooking the metric that reveals how many years your SSD will last under load of your workload. This metric is called write amplification. It is the ratio of physical bytes written to the drive versus the logical bytes your database thinks it sent. It is expressed as the ratio between the logical bytes that your database believes it has sent to the drive, compared to the number of physical bytes that actualy get written on the drive.

For example, if your app sends a single gigabyte of data but the drive winds up sending five then you have a write amplification factor of five. The difference is where endurance die. Use the calculator at the top. Plug in your system’s daily load and type of engine it runs. It will do all the math for you, so you don’t have to guess how much overhead your storage stack adds.

What Is Write Amplification?

Naturaly, most engineers considers only the logical writes; it’s the physical writes that wear out that NAND. Between your silicon and your SQL query is a series of layers of indirection, each adding weight to final count. Let’s begin with the database engine itself. In MySQL or PostgreSQL, for example, the system use a B-tree structure. An update causes indexes to be rebalanced, which leads to page splits that rewrite blocks of unchanged data. LSM-based engines such as Cassandra or RocksDB append your writes and compact them out eventually. They shift the extra work from random I/O during writes to a background merge process. So both mechanisms results in overhead, but at different times and in different ways. And that’s what most folks forget when benchmarking storage engine.

Oh yeah, then there’s the journaling layer. You write ahead some logs called write-ahead logs (WAL). Other people call them binlogs, and others might call them redo logs. I won’t get into that mess here. To put it simply, these things record updates before committing those update to the main table. You’re effectively writing the same data twice. You are writing to both the tablespace and to the log. On top of everything else, if you have a high WAL factor, then you’re effectively doubling your physical write load. Compression will help lessen that to an extent because it makes what you write smaller, but it can never completely remove the underlying duplication caused by need for durability.

Another big amplifier is secondary indexes. Every one of them is its own data structure which needs to be kept up-to-date with main table. What happens when you update or insert into a high churn table? Five spots on disk get touched if you have four secondary index. Reduce your amplification by dropping unused indexes; this is pretty much always the first thing you do. Getting rid of an index you might want some day sounds scary, but maintaining it mean wearing out your hardware and slowing down your queries every day. How fast do you want your query? vs. How fast do you want your hardware to run?

The flash translation layer add its own tax. Because an SSD can’t overwrite cells in-place, it has to erase whole blocks then write new data into them. If there isn’t much free space, the drive will have to move other valid data around to create some free space to accommodate new writes. That garbage collection process increase the write amplification even more. Keeping some amount of free space available allows the drive to work more efficiently and also helps data re-arranging.

This provides a breakdown based off what is contributing to this and allows you to see where your largest leak is. Is it too many indexes? Is it your aggressive compaction settings? Or is it that your SSD just isn’t up to snuff for the workload? Being able to understand the cause of the amplification means being able to fix the correct problem rather than just throwing more drives at it. You can adjust your replication strategy, change the storage engine parameters, or tune the database itself.

The bottom line: Write amplification is a life-vs.-speed-and-safety tradeoff. Turning off index(es) or journal(s) will lower the amplification factor, but slow down queries and/or increase risk of losing data. It’s about controlling write amplification, not removing all overhead. Do so only as far as your hardware allows. Monitor the wear rate daily. When you’re outlasting endurance, examine the breakdown. A minor configuration adjustment may show you have a large surplus in physical writes.

Write Amplification Calculator

Related posts

Leave a Comment