Kafka Cluster Size Calculator

July 12, 2026

Kafka Cluster Size Calculator

Estimate broker count, replicated storage, network throughput, and partition density from daily ingest, retention, replication, compaction, rack awareness, segment overhead, and broker limits.

⚙Kafka workload presets
📊Cluster sizing inputs
Logical compressed data written by producers per day.
Time-based retention target before deletion or compaction cleanup.
Sum of partitions across topics before replicas.
Keep room for OS, logs, leader movement, and emergency retention.
Sustained broker network and disk budget after normal tuning.
1.0 means consumers read roughly the same volume as producers write.
100 for delete retention; 40 means compaction leaves about 40% of bytes.
Index files, time indexes, transaction markers, and small segment slack.
Counts replicas, not just leaders. Lower this for small JVM heaps.
Results are planning estimates; validate with producer batching, ISR behavior, disk class, and workload tests.

Kafka cluster sizing results

Broker Count
3
minimum production footprint
Storage Required
2.7 TB
replicated and buffered
Network Throughput
8 MB/s
producer, replica, consumer
Partitions Per Broker
96
replicas included
Sizing is ready.
🖥Broker spec grid
Mini lab2 TB raw disk
80 MB/s sustained
1 GbE to 2.5 GbE
Small KRaft or dev cluster
SSD standard8 TB raw disk
250 MB/s sustained
10 GbE preferred
General production default
NVMe performance12 TB raw disk
700 MB/s sustained
25 GbE preferred
High ingest and replay
Dense storage24 TB raw disk
220 MB/s sustained
10 GbE minimum
Long retention archives
📚Reference tables
Workload presetTypical inputRetention patternSizing watch point
Home lab events50 to 200 GB/day3 to 14 daysKeep at least 3 brokers for RF 3 testing.
IoT telemetry300 to 1000 GB/day7 to 30 daysMany topics can raise partition metadata load.
CDC pipeline500 to 3000 GB/day3 to 14 daysCompaction reduces storage but not peak ingest.
Audit archive50 to 500 GB/day30 to 365 daysDisk capacity usually dominates broker count.
Kafka factorPlanning valueWhat it affectsPractical limit
Replication factor3 in productionStorage, inter-broker traffic, availabilityBroker count must be at least RF.
Rack awareness2 to 3 racks/zonesFailure domain spreadUse at least as many brokers as active racks.
Disk usable percent70% to 80%Capacity before expansionAvoid running Kafka disks near full.
Segment overhead2% to 8%Log index and slack storageHigher with many small topics or segments.
Partitions per broker1000 to 4000 replicasController, heap, recovery timeTest higher counts before production use.
Network componentFormula usedExample meaningWhy it matters
Producer ingressDaily GB / 86400 secondsAverage write rateProducer spikes require extra headroom.
Replica trafficIngress x (RF - 1)Follower fetch trafficCross-rack traffic can be the real bottleneck.
Consumer egressIngress x read copiesReplay and analytics readsMultiple groups can exceed producer writes.
Total broker budgetIngress + replicas + egressCluster aggregate MB/sBroker count must satisfy throughput and disk.
Cluster sizeGood fitStorage postureOperational note
3 brokersSmall production, home lab HALimited failure and expansion roomRF 3 works, but maintenance is tight.
5 to 6 brokersBalanced team platformBetter leader spread and rolling workCommon starting point for serious workloads.
9 to 12 brokersHigh ingest or mixed consumersMore room for reassignmentsWatch controller metadata and partition count.
18+ brokersLarge platform or long retentionExpansion by rack or storage tierAutomate placement, quotas, and monitoring.
💡Kafka sizing tips
Size twice: calculate once for retained bytes and once for sustained throughput. The larger broker count is the real starting point.
Do not hide consumer traffic: replays, stream processors, and analytics readers can make egress larger than ingest.
Leave disk slack: broker replacement, partition reassignment, and temporary retention spikes all need free space.
Use rack awareness deliberately: RF 3 across 3 racks is strong, but only if leaders and replicas are actually balanced.
Planning note: Kafka sizing is workload-sensitive. Validate with representative message sizes, compression type, acks, batching, consumer lag, tiered storage settings, controller mode, and broker hardware before committing capacity.

The art of sizing a Kafka cluster isn’t as much about raw power as it is about handling constraints. “You’re not buying disks; you’re buying operational sanity, throughput headroom and resilience.” Most engineering teams early on make mistake of sizing strictly for peak ingest at the expense of silent costs of index overhead, replication and consumer replay. Those hidden loads add up and then your brokers will hit disk limits long before they hit network limits.

First, start with retention policy, since it determines your storage footprint. How much can you store in a three-broker cluster? You can only store what will fit if your daily volume is small and you need to keep data for thirty days. It’s not difficult arithmetic, but it’s merciless arithmetic. Your daily ingest times your desired retention period. Times your replication factor. For production, a factor of three is normal: it protects against loss from failure of one broker while preserving availability. That triples your storage requirement right away.

How to Size Your Kafka Cluster

And most folks overlook the fact that this replicated data creates network traffic too. Followers has to acknowledge every write. Inter-broker traffic therefore grows linear with your replication factor. Disk speed isn’t always limiting factor, it’s frequently network throughput. Data is written into Kafka by producers, then read from it by consumers. Typically there are more readers (and they typicaly read more) than writers. Downstream applications may have their own independent offsets; stream processors and analytics engines may play back topics several times over. Your network load will double if your consumer traffic matches your producer traffic, not counting replication traffic across racks. To account for this, the calculator lets you specify headroom percentages and consumer read copies. This can be a helpful way of visualizing how much exit bandwidth you actualy require compared to what you thought was required purely from ingest.

The other silent killer is partition density. The partitions are what give you parallelism, but each one also takes up memory on the broker side. There’s metadata overhead for every single one. You don’t want too many partitions per broker or else the Java virtual machine will spend more time managing its heap space than it does moving data. The rule of thumb here is that you should of keep the number of replicas per broker below four thousand. Want massive parallelism with high throughput? Then you’re going to need more brokers (not necessarily bigger ones). By adding more brokers, you spread out the partition load and make the controller responsible for fewer bits of metadata at a given point in time.

In practice, this is where plans break down. Never let your Kafka disk get full. Never. Leave room for the operating system to rotate segment files, expand index files, and retain data just-in-case of an unexpected traffic spike. Seventy-five percent is plenty of slack. It’s enough for you to rebalance partitions and move leaders around without risking getting hit with a disk-full error partway through. And it provides some wiggle-room for basic maintenance tasks like adding a new node.

There is another wrinkle called rack awareness. With a three-rack setup and a three-replica (replication factor), you can have one replica per rack, where each rack holds one leader and two followers for different sets of partition. You distribute the load across all racks. But with a three-replication-factor and just two racks, one rack ends up with more replicas than the others, leading to uneven loads across both networks and disks. These planning tools here let you test these constraints and avoid building out an uneven cluster which ends up fighting itself at scale.

So, at the end of the day, what is a Kafka cluster? It is a trade-off engine. Trade network bandwidth to get safe replication, brokers for parallelism, and disks for retention. Sizing is more about managing constraints than raw power, there isn’t a perfect size. It is just the right size for today’s set of constraints. Take these as an initial estimate and confirm it by testing out actual latencies in production. Better to slightly over provision day one then rearchitect later when the traffic arrives.

Kafka Cluster Size Calculator

Related posts

Leave a Comment