Hash Collision Probability Calculator
Estimate birthday-bound collision risk for content hashes, IDs, cache keys, trace IDs, backups, and home lab datasets.
Birthday-Bound Results
| Effective bits | About 1 ppm risk | About 1% risk | About 50% risk |
|---|---|---|---|
| 32-bit checksum | 93 hashes | 9.3 thousand | 77 thousand |
| 64-bit fingerprint | 6.1 million | 609 million | 5.06 billion |
| 96-bit truncated digest | 398 billion | 39.9 trillion | 331 trillion |
| 128-bit digest | 26.1 quadrillion | 2.62 quintillion | 21.7 quintillion |
| 160-bit SHA-1 digest space | 1.71e21 hashes | 1.71e23 hashes | 1.42e24 hashes |
| 256-bit SHA-256 space | 1.52e35 hashes | 1.53e38 hashes | 4.01e38 hashes |
| Algorithm or ID type | Typical bits used | Collision use case | Home lab note |
|---|---|---|---|
| CRC32 | 32 | Accidental error check | Too small for unique file identity at scale. |
| Murmur3 32-bit | 32 | Hash table bucketing | Good for distribution, not durable identity. |
| xxHash64 | 64 | Fast non-crypto fingerprint | Usable for small local indexes with verification. |
| UUID v4 | 122 random bits | Distributed object IDs | Strong random ID space when generated correctly. |
| MD5 | 128 | Legacy content fingerprint | Birthday math is large, but adversarial safety is poor. |
| SHA-1 | 160 | Legacy object naming | Avoid for security decisions; collision attacks exist. |
| SHA-256 | 256 | Content addressing | Default conservative choice for backup and dedupe indexes. |
| BLAKE3-256 | 256 | Fast content addressing | Useful for high-throughput file scans and manifests. |
| Dataset size | 64-bit collision risk | 128-bit collision risk | Practical interpretation |
|---|---|---|---|
| 1 million hashes | 0.0000027% | About 1.5e-27 | 64-bit usually fine for temporary local indexing. |
| 100 million hashes | 0.027% | About 1.5e-23 | 64-bit needs a collision check path. |
| 1 billion hashes | 2.7% | About 1.5e-21 | 64-bit is no longer a comfort zone. |
| 1 trillion hashes | Nearly certain | About 0.00015% | 128-bit is still small risk but not magic. |
| 1 quadrillion hashes | Nearly certain | 0.15% | Use 160-bit or 256-bit unless collisions are tolerable. |
| Design | Recommended stored bits | Verification step | Reason |
|---|---|---|---|
| Photo or media NAS catalog | 128 to 256 | Compare file size and full digest | Duplicate libraries can grow for years. |
| Backup manifest | 256 | Store path, size, mtime, and digest | Restore integrity matters more than speed alone. |
| Cache key from URL and headers | 128 | Keep canonical input string | Lets you recheck suspicious collisions. |
| Observability trace ID | 64 to 128 | Partition by time and service | High event counts make short IDs risky. |
| Content-addressed object store | 256 | Verify bytes before dedupe | Hash-only equality is a storage policy decision. |
If you’re like me, you believe that a hash guarantees uniqueness, until it doesn’t. And then you realize it didn’t work when some program you use to back up your computer overwrote one of those different files because they has the same fingerprint. That’s the birthday paradox at play: the math demonstrates the probability of collisions, and shows that these is far more likely than you might think. It is so powerful that it can silently destroy your data if you don’t understand how it works.
People tend to think of hashes as linear buckets. But that’s not how the risk scales, Risk doesn’t work that way at all! Risk doesn’t scale like the number of items; it scales with the square of the number of items. That’s why a 32-bit checksum is safe for storing a couple hundred files, but dangerous for millions.
Why Hashes Can Fail and How to Stay Safe
There’s some crazy math involved, and this page’s calculator does it for you so you don’t have to. But it’s important to know what it’s measuring. You’re defining your total population size $k$ by typing in the batch size and the number of items you already have. Then it compare that to the hash space size $N$, depending on the bit length.
In fact, most will pick their algorithm based off reputation (is it secure?) and performance (how fast is it?). Few consider the effective bit length. If you only store the first 16 characters of a SHA-256 hash in your database, you’ve effectively truncated it to 64 bits. Then you’re storing 64 bits of entropy. Not the full 256 bits of SHA-256. For the use of the calculator, that’s what we accept as the actualy truth. That’s what most systems fail at: truncation.
Using a new crypto function is good. But shortening that to something smaller for old school schema increase the chance of error. What about a distributed cache versus a local photo library? A collision in a distributed cache may return the incorrect asset to a user, which breaks functionality. But it’s an inconvenience if a collision occur on a local photo library; you can use metadata like file size to reveal the collision before data is corrupted.
The tool lets you adjust the number of namespaces so you can simulate sharding. This is important: data sharded into separate partitions has reset the collision risk for each partition. Independent shards mean less risk. Global index shards with labels mean there is a cumulative risk.
Fortunately the tool comes with some sanity checking reference tables. At about five billion items there’s a 50% chance of a 64-bit fingerprint hitting an existing value. That may sound like a lot, but it’s well within reach of high-frequency logging and IoT telemetry. If your system has to run years, 128 bits or greater is the normal comfortablty zone.
Many criticize MD5 for being cryptographically weak. It certainly isn’t good for adversarial collision resistance, but as far as non-adversarial collision resistance goes, its 128-bit space are still statistically big enough for most hobbyist project. Where you get into trouble is assuming full strength where you are really only using a fraction of the output. But it needs to be verified too.
In a limited space, there’s no such thing as zero risk. When a hash is used, consider it merely an identifier (at best) unless backed up by some other checks. Save the hash along with some data about that file: its path or size or even its contents. A collision will then show itself before any data has been harmed through a simple check of those metadata fields.
It costs nothing to do this defensively. To not do so and later regret it could of been costly. To pick a hash length you must strike a balance between safety margin vs storage efficiency. How many additional bits do you want to allocate to having peace of mind? The numbers is in the calculator. Pick your tolerance for failure and go from there.
After seeing how fast the probability goes up, you’ll probably opt for longer hashes which use more disk space. Better to waste some bytes than not be able to tell one file apart from another.



