Erasure Coding Overhead Calculator

September 11, 2026

HomeServerBlog distributed storage planner

Erasure Coding Overhead Calculator

Estimate usable capacity, parity overhead, safe fill, logical capacity after compression, fault tolerance, rebuild read load, and rebuild time for k+m erasure coded home lab storage pools.

▦Erasure coding presets

⚙Erasure coding inputs

Drive vendors use decimal TB; many filesystems report binary TiB.
k data shards plus m parity shards per stripe group.
Used for rebuild throughput, efficiency, and endurance hints.
Only complete k+m groups contribute usable capacity.
Use the smallest drive size if mixed drives are in the same coding group.
More data shards improve raw efficiency but increase rebuild fan-in.
m is the number of shard or drive failures a group can tolerate.
Distributed stores usually need free room for backfill, compaction, and rebalance.
Metadata, checksums, allocation maps, block headers, and object manifests.
Free space kept outside normal fill for recovery and rebalancing.
Use 1.00 for media, encrypted backups, or already compressed data.
Applies after filesystem overhead and rebuild reserve.
Safe Logical Capacity 0 TB after reserves and compression Formula: usable x fill x reductions x compression.
Parity Overhead 0% relative to data capacity Formula: m / k.
Fault Tolerance 0 failures per group Survives up to m missing shards in each group.
Rebuild Estimate 0 hr single lost shard Formula: drive size / effective rebuild rate.

Capacity breakdown

Health and rebuild reading

Ready.

▣Calculated capacity markers

0 TBRaw pool

Total participating drive capacity before parity or reserves.

0 TBEC usable

Complete groups times k data shards times per-drive capacity.

0 TBParity stored

Complete groups times m parity shards times per-drive capacity.

0Unused drives

Drives outside complete k+m groups unless your platform spans groups differently.

⚒Equipment and media comparison grid

NAS HDD 5400/5900

95 MB/s

Quiet capacity drives with lower rebuild speed. Useful for media pools and small always-on storage.

NAS HDD 7200

145 MB/s

Balanced home lab choice. Higher sequential rebuild rate but more heat and vibration than slower disks.

Enterprise HDD

185 MB/s

Better sustained rebuild behavior and workload rating. Watch idle watts in a small rack.

SMR Archive HDD

45 MB/s

Capacity-focused media can rebuild slowly during random rewrites, compaction, or backfill.

SATA SSD

420 MB/s

Predictable small-block recovery with SATA bus limits. Endurance reserve matters for write-heavy pools.

NVMe SSD

1200 MB/s

Fast rebuild math, often limited by network, CPU coding work, or PCIe topology.

USB Backup HDD

75 MB/s

Good for portable vaults, less ideal for always-on degraded recovery due to bus and enclosure limits.

Mixed Used Drives

65 MB/s

Model using the smallest and slowest member when old disks share one stripe group.

ℹReference tables

Capacity by k+m layout

LayoutRaw efficiencyParity overheadBest use
3+260.0%66.7% over dataSmall clusters, high protection per node.
4+266.7%50.0% over dataCommon six-drive or six-node home lab groups.
6+275.0%33.3% over dataGood balance for eight-drive NAS shelves.
8+280.0%25.0% over dataCapacity-heavy media with reliable hardware.
8+372.7%37.5% over dataLarge disks where rebuild exposure matters.
10+471.4%40.0% over dataWide archive groups with four-shard tolerance.

Drive capacity conversions

Drive labelApprox TiB4+2 usable6+2 usable
4 TB3.64 TiB16 TB / 14.6 TiB24 TB / 21.8 TiB
8 TB7.28 TiB32 TB / 29.1 TiB48 TB / 43.7 TiB
12 TB10.91 TiB48 TB / 43.7 TiB72 TB / 65.5 TiB
18 TB16.37 TiB72 TB / 65.5 TiB108 TB / 98.2 TiB
22 TB20.01 TiB88 TB / 80.0 TiB132 TB / 120 TiB
30 TB27.28 TiB120 TB / 109 TiB180 TB / 164 TiB

Rebuild and backfill planning

Media profileModel rateEfficiencyPlanning note
NAS HDD 5400/590095 MB/s70%Quiet pools need longer degraded windows.
NAS HDD 7200145 MB/s75%Balance rate against heat and vibration.
Enterprise HDD185 MB/s80%Better for heavy scrub and recovery cycles.
SMR Archive HDD45 MB/s55%Avoid heavy degraded random write workloads.
SATA SSD420 MB/s82%Often limited by controller or network path.
NVMe SSD1200 MB/s75%CPU coding or network fan-in can dominate.

Common project sizes

ProjectLayoutRaw poolSafe result
Six-bay mini cluster4+248 TBAbout 21 to 24 TB logical.
Eight-bay home NAS6+296 TBAbout 48 to 56 TB logical.
Ten-disk media shelf8+2180 TBAbout 105 to 120 TB logical.
Triple-parity photo lab8+3220 TBAbout 115 to 130 TB logical.
Dual archive groups10+4504 TBAbout 260 to 300 TB logical.
Small edge nodes3+220 TBAbout 8 to 10 TB logical.

⚡Home lab planning tips

Keep the fault domain honest: A 6+2 group protects two missing shards in that group, not two whole racks unless shards are deliberately placed across rack or node failure domains.
Capacity is not the only overhead: Wide k values look efficient, but every rebuild reads from more surviving shards and can stress networks, OSDs, CPUs, and small write paths.

This calculator models erasure coding at the planning layer. Real platforms may reserve additional space for placement groups, object size rounding, journals, snapshots, compression dictionaries, checksums, or minimum allocation units.

After you’ve got a bunch of disks stuffed inside a server chassis, how many are left over? And what’s the server holding back as part of the file system protection?

If you use erasure coding, it’ll break up your data into shards that get spreaded out onto different disks. So if some goes missing, you won’t lose your files. But there’s no free lunch here. Rebuilding from scratch takes time, and there’s an overhead cost in terms of matching. Raw disk capacity comes with a price tag, too.

How to Use the Erasure Coding Calculator

Plug in your number of drives and type of coding on top and the calculator will work the numbers out for you… No need to guess at what conversion factor or coefficient applies.

K plus m is the core idea here. K equals the data shards containing your real data. M equals the parity shards needed for rebuilding when a drive go down. A typical home lab deployment will use four drives for data and two for parity. That’s sixty-six percent efficiency, i.e., two-thirds of available disk space. But amount of parity data used to protect information is 50% larger.

By moving to a six plus two layout, you achieve seventy-five percent efficiency. You get more bang for your buck on parity. But more expense during rebuild: Each time a drive die in a six plus two stripe set, the system needs to draw from five surviving data drives to rebuild. It is efficient at rest but expensive in motion.

But also take into account what kind of media it’s going to be. A fast seven-thousand-two-hundred RPM enterprise disk will rebuild far faster than a slower five-thousand four-hundred RPM NAS drive. These throughput numbers can vary widely. According to the table on the page, a mechanical drive could rebuild as slow as ninety-five megabytes per second, while an NVMe SSD can reach one thousand two hundred.

Why does this matter? Because the longer your rebuild takes, the higher your chance of having another drive fail before the initial one has been fixed. If you have a big parity group rebuilding over a long time and two drives goes down in the process, you’re out of luck. This is why calculator asks for both a degraded rebuild reserve and target maximum fill percentage.

I’m not being conservative when I leave my pool 80% full. This leaves room for storage engine to breathe. To do garbage collection of SSDs. It handles metadata updates. It also performs backfill operations where data are moved around in distributed systems. If you pack every terabyte without leaving any space, the system slows down as it looks for empty slots where it can write its new parity information.

Another strong way is adding a compression ratio. If you’re storing un-compressed video, databases or text logs, then you may get a ten percent gain. But if you are already compressing media files or want to encrypt backups, the ratio doesn’t change. It stays one to one, which means capacity is actualy what the math says.

What about layout? You will want different layouts depending on the type of failure you expect in your fault domain. If you’re storing a media archive and don’t change much, you can get four-shard tolerance and huge capacity by going 10 plus four. It takes ten drives, but they all function as a single unit.

If space isn’t an issue, then go with safety: 3 plus 2 protects against two drives in a very small subset failing at once. If you have a busy home server that does lots of writes, go with a lower k, it’ll keep the write penalty down. If you’re doing more long-haul stuff, go with a higher m and rest easy knowing that you’ve got some peace of mind.

The tool also figures out how long it would of take to rebuild one shard, so you can kind of visualize how long that’s taking. That makes a difference when planning your maintenance windows, too. It could be a six-hour rebuild instead of a 12-hour rebuild.

In conclusion, erasure coding means sacrificing some space to make it safer. Storage is not free. You can’t get maximum speed + maximum safety + maximum density, you’ve got to choose two. Play with the presets; they’ll show you what impact various profiles have on the numbers. Try a deep archive vault versus a mini Ceph cluster, and look at how the usable capacity scales when the parity changes in size.

The point here is to configure a pool that withstands expected hardware failure while not forcing you to buy more disk space than you wanted. Begin with real-world hardware inputs, tweak the fill levels based off comfort, and then let the numbers inform your buying decisions. Your data will appreciate the foresight.

Erasure Coding Overhead Calculator

Related posts

Leave a Comment