Ceph Erasure Coding Calculator for Pools

July 9, 2026

Ceph Erasure Coding Calculator

Estimate Ceph EC k/m pool efficiency, raw and usable capacity, PG count hint, failure-domain fit, object padding, and recovery read amplification.

⚙ Ceph pool presets
🗄 EC profile and cluster inputs
Logical data shards in the erasure-code stripe.
Parity shards; normally equals tolerated shard failures.
The calculator checks whether k+m chunks can spread cleanly.
Hosts, racks, or domains available for placement.
All OSDs that may receive this EC pool.
Decimal TB per OSD after device selection.
User data before EC overhead and reserve.
Leaves room below full ratios and during recovery.
Used only for a planning hint; check Ceph autoscaler.
Scales the PG hint for multi-pool clusters.
RADOS object size for padding and stripe estimates.
Approximate EC profile stripe unit for object padding.
Usable EC Capacity
0 TB
after reserve
Raw Required
0 TB
for planned data
PG Count Hint
0
rounded to power of 2
Recovery Read Amp
0:1
per missing shard
Run the calculator to see placement and capacity notes.
🧮 Ceph EC profile grid
4
Data chunks k
2
Coding chunks m
6
Acting set width
66.7%
Storage efficiency
📊 Reference tables
EC profile Width k+m Raw overhead Failure tolerance Typical home lab use
2+13 OSDs minimum1.50x raw1 shardSmall test pool, not much spread
3+25 OSDs minimum1.67x raw2 shardsCompact cluster with extra parity
4+26 OSDs minimum1.50x raw2 shardsGeneral NAS data and backups
6+28 OSDs minimum1.33x raw2 shardsBalanced media and VM images
6+39 OSDs minimum1.50x raw3 shardsArchive pool with more fault room
8+311 OSDs minimum1.38x raw3 shardsLarge rack-aware object pool
10+414 OSDs minimum1.40x raw4 shardsCold object data on wide clusters
PG planning item Formula used here Why it matters Practical note
Base PGsOSDs x target x share / widthEC PGs touch k+m OSDsRound to a power of 2
Acting set widthk + mEach PG maps that many shardsNeeds enough distinct OSDs
Autoscaler checkUse final cluster statsBalances pools by bytesPrefer Ceph mgr advice
Too few PGsHot OSD riskLess placement spreadWatch OSD fullness skew
Too many PGsMemory and peering loadMore metadata workIncrease carefully
Object size 4+2 stored 8+3 stored Padding sensitivity Best fit
64 KiB6 stripe units11 stripe unitsHighPrefer replicated metadata
1 MiBAbout 1.5 MiBAbout 1.4 MiBMediumSmall object buckets
4 MiBAbout 6 MiBAbout 5.5 MiBLowRBD or object data
16 MiBAbout 24 MiBAbout 22 MiBLowBulk media and archive
Failure domain Input meaning Minimum useful count Risk if short Planning action
OSDEach shard on a diskk+m OSDsHost loss may lose many shardsUse only in tiny labs
HostEach shard on a nodek+m hostsUndersized acting setsAdd nodes or lower width
ChassisEach shard per chassisk+m chassisChassis fault can exceed mUse device classes too
RackEach shard per rackk+m racksPlacement may be impossibleUse smaller k/m profile
RoomEach shard per roomk+m roomsStretch design pressureConfirm CRUSH map first
💡 Ceph planning tips
Failure domain tip: For host, rack, or room domains, the cleanest EC layouts have at least k+m domains available so CRUSH can place one shard per domain.
Recovery tip: EC pools save raw capacity, but rebuilding a missing shard must read k surviving shards. Wide profiles improve efficiency but increase degraded read work.
This calculator gives planning numbers for Ceph erasure-coded data pools. Validate final PG counts, crush-failure-domain rules, device classes, min_size, and autoscaler output on the actual cluster before creating production pools.

The math said you needed a certain amount of raw storage across your hard drives to store your media library and virtual machines securely. You purchased those drives. Then you configured an erasure coding pool and saw that your usable capacity was significantly below what you expected. If this sounds familiar, then you’ve transitioned from a basic replication pool to a Ceph erasure coding pool and are wondering how this could be.

It’s not. It’s simply arithmetic; an exchange between faster recovery time and more efficient usage. Before you initialize your very first disk, understand that tradeoff so that you don’t build out a cluster which might look great on paper, but doesn’t perform well in reality.

Why Your Storage Space is Smaller Than Expected

Plugging in your data and desired level of parity turns over the work to the calculator, which will do the math for you. The calculator handles the math once you plug in your data and parity requirements, which removes the guesswork from coefficients and conversions. With a 4+2 profile, it’s telling you that your storage efficiency is about sixty-seven percent; four parts are actualy data, while two are parity.

Sounds good until you realize this means if you lose a single shard, the system has to read all four others and try to put together what was lost. The larger your stripe, the more network traffic you’re going to create as things rebuild themself. It is small stuff but it is important when you want to get your data back fast without clogging up your network bandwidth.

Where most home lab admins go wrong is choosing the correct failure domain. Whatever CRUSH map you create, it should spreads its shards over hosts, so make sure you have enough physical machine to match that setting. For a profile of 4+2, for instance, you’ll need at least six different hosts (which means one shard per host). Anything less (such as four hosts) and your pool’s placement rules will conflict and your pool won’t be healthy.

It doesn’t matter how much you fiddle with software; you can’t fake geography. Your hardware has to back up whatever geographic distribution strategy you choose in your settings. You should also consider object size, which isn’t shown in capacity charts. With erasure coding, the idea is to split data into stripes, and smaller objects may be padded out so they take up full stripe. For example, if a wide EC pool has stripes that require padding, a little text file will occupy as much on-disk space as big video file, since it’s padded out to the next stripe boundary. That’s one reason why erasure coding tends to make more sense for bulk storage, as opposed to a workload involving many small but metadata-rich objects. Replication may actualy result in fewer wasted bytes (and higher performance) compared to erasure coding if your workload consists of millions of small files.

The other thing that takes some forethought is how you plan out your placement groups. If you have too few, you’ll end up with hot spots when some drives fill up before others because they’re unevenly distributed among OSDs. Having too many PGs makes cluster peering tasks slower and also increases amount of memory the monitors use. Again, there’s a hint provided by the tool based off your target load and number of OSDs but you should confirm it against the Ceph autoscaler after the cluster is live. It doesn’t give you the final decision; it gives you a place from which to tune.

The price of efficiency you don’t see is that your recovery gets amplified. If a drive dies on a pool where there are replicas, it just copies data from another drive. With an EC pool, it has to pull out chunks from several survivors and stitch them together to recreate the parity. So with a 10+4 profile, saving a lot of disk space versus replicating, if a shard dies, then you have to read in ten others to rebuild it. It causes a lot of I/O on those other drives while you’re trying to recover. You pay less operationally when things are fine; you just get more capacity, but you would of paid more operatively when something goes down.

Your goal isn’t perfection, it’s matching your storage approach to your true tolerance for complexity and risk. Wider profiles work if you’ve got lots of drives and want maximum durability. Simpler ones will help you stay in control if your CPUs is weak or your network bandwidth is limited. Once a pool exists, changing its profile is hard; that’s why getting the starting layout correct helps ensure lasting growth (not a maintenance headache) as your cluster scales. Go with what makes sense to you, check the capacity implications, and use the numbers to inform your hardware buying decisions, not guesswork.

Ceph Erasure Coding Calculator for Pools

Related posts

Leave a Comment