Lambda Concurrency Calculator
Estimate required AWS Lambda concurrency from requests per second, average and p95 duration, burst multipliers, reserved concurrency, provisioned concurrency, retries, async backlog, and account limits.
Steady synchronous or event arrival rate.
Mean Lambda execution duration in milliseconds.
Use p95 to model burst occupancy and tail latency.
Short-term traffic multiple above steady RPS.
Function-level reserved concurrency. Use 0 for none.
Warm execution environments allocated ahead of time.
Extra invocations from client, service, or async retries.
Queued events to drain alongside fresh traffic.
Target time to clear the async backlog.
Regional account concurrency quota available to Lambda.
Headroom for jitter, regional variance, and uneven traffic.
Changes the risk message and backlog interpretation.
Retry load expands invocation demand before concurrency math.
Async queued events converted to a drain rate.
Function-level ceiling when reserved concurrency is set.
Regional concurrency limit shared by all functions.
| Metric | Formula | Use In Calculator | Operational Note |
|---|---|---|---|
| Steady concurrency | RPS x avg seconds | Base required concurrency | Best for normal API traffic and warm sizing. |
| Retry-adjusted load | RPS x (1 + retry rate) | Expands fresh invocation demand | Retries can double pressure during partial failure. |
| Backlog drain rate | backlog / drain seconds | Adds async catch-up demand | Use for SQS, EventBridge, and async invokes. |
| Burst concurrency | RPS x burst x p95 seconds | Tail-aware spike estimate | p95 protects against long-running overlap. |
| Provisioned gap | required - provisioned | Shows cold-start exposure | Provision steady hot path first, not rare spikes. |
| Workload Pattern | Typical Duration | Concurrency Driver | Reserved Concurrency Hint | Provisioned Concurrency Hint |
|---|---|---|---|---|
| API Gateway request handler | 50 to 300 ms | RPS and p95 latency | Set above burst concurrency if critical. | Warm the steady p50 to p90 path. |
| Webhook receiver | 100 to 800 ms | Vendor fan-in and retry waves | Protect downstream services with a ceiling. | Useful for predictable campaign windows. |
| SQS queue worker | 500 ms to 5 sec | Backlog drain target | Cap concurrency to match database capacity. | Often less important than batch tuning. |
| Stream processor | 100 ms to 2 sec | Shard count and batch size | Coordinate with shard parallelization. | Warm only strict low-latency processors. |
| ML or image inference | 1 to 30 sec | Long duration overlap | Use reserved caps to avoid account starvation. | Provision if cold starts hurt SLOs. |
| Runtime | Cold Start Profile | Duration Tendency | Concurrency Planning Note |
|---|---|---|---|
| Node.js | Low to medium | Fast for I/O APIs | Provision only latency-critical paths. |
| Python | Low to medium | Good for glue and data tasks | Watch package size and import time. |
| Java | Medium to high | Strong after warm-up | Provisioned concurrency often pays off for APIs. |
| .NET | Medium to high | Predictable business logic | Use p95 duration when cold starts are visible. |
| Go | Low | Efficient CPU and network tasks | Often needs less provisioned headroom. |
| Container image | Depends on image size | Workload-specific | Keep image lean before buying more warm capacity. |
| Preset | RPS | Avg / p95 | Burst | Why It Fits |
|---|---|---|---|---|
| Home API Gateway | 120 | 90 / 180 ms | 1.8x | Small public API with modest tail latency. |
| Webhook Fan-In | 700 | 160 / 650 ms | 4x | Vendor retries and delivery waves arrive together. |
| SQS Worker | 60 | 1200 / 2600 ms | 1.4x | Backlog drain dominates the concurrency target. |
| ML Inference | 25 | 4500 / 9000 ms | 2x | Long overlap makes concurrency climb quickly. |
| Public Launch Spike | 1800 | 140 / 500 ms | 5x | Short requests still need large burst allowance. |
In the past, trying to handle sudden bursts of serverless traffic felt like playing the lottery. You launch that cool new feature, see your log files start filling up with throttling errors, and rush to increase your limits before your users complain. It’s been a game of guessing and hoping you’re right.
The underlying issue isn’t so much “does Lambda work well” under ordinary circumstances, but what happens when reality no longer resembles your test environment. How many concurrent executions will likely occurs at the same time? It is not about number of requests per second. On paper, it seems straightforward: the basic math. Steady state concurrency is (requests per second) x (average execution duration).
Why Your Serverless App Fails During Traffic Bursts
But very few production systems spends their time in this neat little steady state. Production systems have backlogs from previous outages that is still eking their way through the system. They have tails and they have retries. And most estimates break down there. Plan for the average. Then one day when you see a burst, you’ll hit the wall because you didn’t account for all those slow ones that built up on top of each other.
This is easier to visualize if we use the calculator above. It will take your baseline traffic and multiply that by your p95 duration, which is how long those requests last. How long those requests last), which blows up your concurrency requirements on short bursts. Then it adds back async backlog draining rates + retries to show the overall load being thrown at your functions. It doesn’t spit out just “steady state” number, either. It breaks apart the burst requirement vs background noise, and shows you where the bottlenecks are forming… Because the way you prevent a bottleneck is different than the way you fix one.
Retries are the silent killer of concurrency planning. It’s tempting to think “A few failed requests won’t hurt”. Then you realize they’re automatically retried and now you’ve doubled (or tripled) your load within a few seconds. That’s a multiplier effect, something the calculator takes into account so you can visualize how much additional capacity you’ll need just to survive a minor hiccup from one of your downstream systems. Even a healthy function can throttle itself if its retry rate is too aggressive. Resilience vs. Resource exhaustion: you gotta balance these two things.
And then there’s asynchronous backlogs. This means when your events stack up in EventBridge and SQS, they don’t dissapears. Instead, they sit in a queue awaiting some available room to process them, but that same room is required to process new traffic too. To process all those backlog event, it needs dedicated concurrency; concurrency that could otherwise serve real-time requests. Unless you account for the rate at which the backlog drains, your system will either throttle new users attempting to reach the front of the queue, or it’ll hang around processing old stuff forever.
The calculator assumes this backlog is actualy traffic demand, so you need to decide how fast you want to empty the queue vs. You must decide how much spare capacity to reserve for spike periods. The levers you have to manage this are reserved and provisioned concurrency. Reserved concurrency provides a hard ceiling: it won’t let one runaway handler starve other functions. Provisioned concurrency eliminates cold starts by keeping your environments warm. This is critical for latency sensitive paths, but it is very expensive if applied indiscriminately.
The tool’s reference tables gives clues as to how they interact across different types of workloads such as batch ETL jobs vs API gateways. While an authentication callback can’t afford a few extra milliseconds of cold start time, an image processing job probably could. The calculation is also different depending on what runtime you choose: In certain configurations, languages like Go and Node.js are generally more adept at handling warm-up vs. This affects how much starting capacity you must set up to provide smooth service.
The less well your language starts up, the greater the cost of getting concurrency wrong, each cold start consumes a chunk of available slots, resulting in longer wait times for answers. The takeaway: Properly sizing Lambda is not about optimizing for the average. It is about managing variance by getting your headroom just right. You need room to spare for those inevitable bursts, but you should of have no extra cost when things are quiet.
The calculator above does the heavy math for you, while you figure out the right retry strategy and set some smart throttle limits. Once you get your head around how backlogs, retries and duration compound each other, throttling isn’t such a black art anymore; it’s just another engineering problem that’s solvable.



