passdrill

Lambda Reserved vs Provisioned vs Unreserved Concurrency: One Traffic Burst, Worked Example

AWS Lambda gives you three separate concurrency levers: reserved concurrency, provisioned concurrency, and the shared unreserved pool everything else draws from. Most explanations cover each one in isolation, or show a single function's timeline in the abstract. This page sets up one small AWS account with three real functions, sends it one traffic burst, and traces exactly which requests get a warm start, which get a cold start, and which get throttled — and why.

The account, before any traffic arrives

By default, every AWS account gets a concurrency limit of 1,000 execution environments per Region, shared across every function. Reserving concurrency for one function carves that amount permanently out of the shared pool: AWS always keeps at least 100 units of unreserved concurrency free for functions that don't reserve any, so the most you can ever allocate across all your reservations is 900.

Here is the account, configured for three functions:

FunctionReserved concurrencyProvisioned concurrencyWhere its capacity comes from
checkout (synchronous, called by the storefront)300200 (out of its own 300)Its own sealed-off 300 slots
nightly-report (scheduled batch job)500Its own sealed-off 50 slots
webhook-ingest (no reservation configured)——The shared unreserved pool

The account's unreserved pool is what's left after both reservations: 1,000 − 300 − 50 = 650, comfortably above the 100-unit floor AWS enforces. If you tried to push checkout's reservation up to 950 instead, the console would reject it outright, because 950 + 50 = 1,000 leaves only 0 unreserved units — below the mandatory 100. This is a real, enforced limit, not a soft guideline: you cannot reserve away every last unit of shared capacity, no matter how many functions you have.

checkout also has 200 of its 300 reserved slots pre-initialized as provisioned concurrency, so 200 execution environments are sitting warm, dependencies loaded, ready to run the handler the instant a request lands. The remaining 100 reserved slots exist as a ceiling and a guarantee, but nothing is pre-warmed there yet.

The burst: 700 concurrent requests hit checkout in under a second

A flash sale starts. checkout jumps from near-zero traffic to 700 concurrent synchronous invocations almost instantly. Here is what happens to those 700 requests, in order:

RequestsWhat happensWhy
1–200Run immediately, no cold startThey land on the 200 pre-initialized provisioned-concurrency environments — nothing to initialize, the handler runs right away.
201–300Run, but each pays a cold startProvisioned concurrency is exhausted, but checkout still has headroom inside its own 300-unit reserved ceiling, so Lambda spins up 100 fresh on-demand environments within that reservation.
301–700Throttled immediately (a synchronous caller gets a 429 TooManyRequestsException)Reserved concurrency is a hard ceiling as well as a guarantee: once all 300 reserved slots are busy, checkout cannot borrow from the 650-unit unreserved pool, even though that capacity is sitting idle two rows down in the account.

That last row is the part that surprises people moving from EC2/ASG thinking: reserved concurrency doesn't just guarantee a floor, it also seals the function off from ever using more than its own reservation, permanently, regardless of how much spare capacity the rest of the account has.

Meanwhile, the other two functions are unaffected — or aren't

At the same moment, nightly-report's scheduled trigger fires and needs 60 concurrent executions, but it only reserved 50. The extra 10 throttle too, completely independently of checkout's flash sale — its 50-unit reservation is just as sealed off in the other direction, and it can't dip into checkout's spare reserved capacity (there isn't any right now) or the unreserved pool either.

webhook-ingest, which never reserved anything, is running a steady 40 concurrent invocations from the unreserved pool. It notices nothing. The 650-unit unreserved pool never overlaps with either reservation, so a spike hammering checkout and a scheduling collision hitting nightly-report both leave webhook-ingest completely untouched. This mutual isolation — not raw capacity — is the actual reason teams reserve concurrency for critical functions: a runaway or misbehaving function can never starve another reserved function of its guaranteed slots.

A second ceiling that has nothing to do with reserved concurrency

Suppose checkout's reservation had been set to 1,000 instead of 300, so none of those 700 requests would ever hit a reserved-concurrency wall. Even then, the burst wouldn't complete cleanly in under a second. Lambda enforces a separate, per-function concurrency scaling rate: at most 1,000 new execution environment instances every 10 seconds, for that function, regardless of how much account-level or reserved-concurrency headroom exists. A jump from near-zero to 700 environments in under two seconds outpaces that ramp, so some of those requests would still queue behind the scaling rate for a few seconds — a throttling reason entirely separate from running out of reserved slots or account-wide capacity. Reserved concurrency, the account limit, and the scaling rate are three independent quotas, and a real burst can hit any one of them first depending on how the numbers line up.

The three levers, side by side

Reserved concurrencyProvisioned concurrencyUnreserved pool
What it isA dedicated floor and ceiling of concurrent executions for one functionA number of environments pre-initialized ahead of any requestThe account's shared leftover capacity
Removes cold starts?No — new environments still initialize on demandYes, up to the provisioned numberNo
Can it be starved by other functions?No, it's sealed offNo, it's part of the reservationYes — it's first-come, first-served across every unreserved function
Extra cost?NoneYes, billed continuously while allocatedNone
What happens when it's exhausted?Throttling (immediate error for sync calls)Falls back to reserved or unreserved concurrency, with a possible cold startThrottling once the account limit is reached

Reserving concurrency answers "how do I stop one function from starving another, or from overwhelming a downstream database?" Provisioning concurrency answers a completely different question, "how do I avoid cold starts for latency-sensitive traffic?" — and you can use either one alone, both together, or neither. Because provisioned concurrency can never exceed a function's reserved concurrency once one is configured, sizing both together (as checkout does above: 200 provisioned inside a 300 reservation) is what lets a function absorb its normal peak with zero cold starts while still tolerating a short burst above that peak without an immediate hard throttle.

For more on how Lambda scales a single function's concurrency in isolation, and how event source mappings and asynchronous invocations behave differently under throttling, see the AWS SAA serverless practice questions, which cover reserved concurrency's throttling behavior, provisioned concurrency's cold-start elimination, and the per-function scaling rate as three separate tested facts.

Source: AWS Lambda Developer Guide, "Understanding Lambda function scaling" and "Configuring provisioned concurrency" (docs.aws.amazon.com/lambda/latest/dg/), retrieved 2026-09-30.

Drill Lambda & Serverless practice questions →