AWS Lambda Costs: How Millisecond Billing Works and What to Tune
Lambda is cheap for spiky and low-volume work and can get expensive for steady, high-volume work. Unexpected bills usually come from three places: more invocations than anyone expected, memory settings nobody tuned, and the services around Lambda that bill on their own.
What you pay for
The Lambda bill has two main parts:
- Requests. A fixed price per million invocations.
- Duration. Billed in GB-seconds: configured memory times execution time, rounded up to the nearest millisecond.
Details that change the math:
- Memory is configurable from 128 MB to 10,240 MB, and CPU is allocated in proportion to memory. At 1,769 MB a function gets the equivalent of one vCPU.
- Functions on
arm64(Graviton) have a lower price per GB-second thanx86_64. - Since August 2025 the INIT phase (cold start initialization) is billed for all functions, including zip-packaged functions on managed runtimes, which used to get it for free. Heavy startup code now shows up on the bill, not only in latency.
- Ephemeral storage above the default 512 MB is billed separately.
- Provisioned concurrency is billed for the time it is enabled, whether requests arrive or not.
Use the AWS pricing page and calculator for current rates in your region. Cost per function is roughly: invocations × (request price + billed duration × memory × duration price). There are two levers: fewer invocations, and fewer GB-milliseconds per invocation.
Measure before tuning
Each invocation writes a REPORT line with duration, billed duration, configured memory, max memory used and, on cold starts, init duration. CloudWatch Logs Insights can summarize it:
filter @type = "REPORT"
| stats count() as invocations,
avg(@billedDuration) as avg_billed_ms,
pct(@billedDuration, 99) as p99_billed_ms,
max(@maxMemoryUsed / 1000 / 1000) as max_used_mb,
max(@memorySize / 1000 / 1000) as configured_mb
by bin(1d)
Add and ispresent(@initDuration) to the filter to count cold starts. Run it per function, then sort functions by invocations × average billed duration × memory. A handful of functions usually account for most of the spend.
Right-size memory
More memory means more CPU. For CPU-bound code, doubling memory can more than halve duration, which makes the invocation cheaper and faster. For I/O-bound code that mostly waits on network calls, extra memory just costs more.
Don't guess. AWS Lambda Power Tuning (an open-source Step Functions state machine) runs a function at several memory sizes and plots cost against duration. Compute Optimizer also gives memory recommendations for Lambda once it has enough history.
Cut invocations you don't need
The cheapest invocation is the one that never happens.
Filter at the event source. A handler that checks the event and returns early still pays for the request and the duration. Event source mappings for SQS, Kinesis, DynamoDB Streams and Kafka support filter criteria, so non-matching records never invoke the function.
Batch. For SQS, raise the batch size and set a batching window so one invocation handles many messages.
resource "aws_lambda_event_source_mapping" "orders" {
event_source_arn = aws_sqs_queue.orders.arn
function_name = aws_lambda_function.orders.arn
batch_size = 100
maximum_batching_window_in_seconds = 5
function_response_types = ["ReportBatchItemFailures"]
filter_criteria {
filter {
pattern = jsonencode({ body = { status = ["created"] } })
}
}
}
For SQS, messages that don't match the filter are deleted from the queue, so only filter out what nobody else needs. With batches, report partial failures (ReportBatchItemFailures) so one bad message doesn't make the whole batch retry.
Avoid synchronous function chains. When function A calls function B and waits, you pay for both durations at the same time. Use Step Functions, an SQS queue, or direct service integrations instead.
Cold starts
- Keep deployment packages small. Bundle and tree-shake (esbuild for Node.js), drop unused dependencies.
- Initialize SDK clients and connections outside the handler so warm invocations reuse them, but don't load what most invocations never need.
- SnapStart restores from a snapshot of the initialized environment. It supports Java 11+, Python 3.12+ and .NET 8+. For Python and .NET there are separate caching and restore charges.
- Provisioned concurrency removes cold starts for the configured number of environments, and you pay for it all the time. Use it for latency-sensitive paths with predictable traffic, and scale it on a schedule.
The costs around Lambda
These often exceed the function cost itself:
- CloudWatch Logs. Ingestion is billed per GB. Set retention on every log group and use log levels. With
log_format = "JSON", Lambda can filter application and system logs by level. - API Gateway. HTTP APIs are cheaper than REST APIs and cover most use cases.
- NAT gateway. Functions in a VPC that call AWS APIs or the internet pay NAT processing per GB. Use VPC endpoints where possible.
- Downstream services. DynamoDB, Step Functions transitions, X-Ray traces.
A baseline function definition:
resource "aws_cloudwatch_log_group" "orders" {
name = "/aws/lambda/orders"
retention_in_days = 14
}
resource "aws_lambda_function" "orders" {
function_name = "orders"
role = aws_iam_role.orders.arn
filename = "build/orders.zip"
handler = "index.handler"
runtime = "nodejs22.x"
architectures = ["arm64"]
memory_size = 512
timeout = 10
environment {
variables = {
TABLE_NAME = aws_dynamodb_table.orders.name
}
}
logging_config {
log_format = "JSON"
application_log_level = "INFO"
system_log_level = "WARN"
}
depends_on = [aws_cloudwatch_log_group.orders]
}
Guardrails
- Reserved concurrency caps how many copies of a function can run, which also caps how fast it can spend.
- Timeouts close to real p99 duration. A hung call bills until the timeout.
- Recursive loops. Lambda detects and stops some loops through SQS, SNS and S3, but design so a function never triggers itself (for example, write output to a different bucket or prefix).
- Budgets and Cost Anomaly Detection on the account or per tag.
When Lambda stops being cheap
For steady, high-throughput workloads, compare against containers on Fargate, ECS or EKS. Use your own measured numbers: requests per second, average billed duration, memory. If the function is busy most of the time, an always-on container is often cheaper.
Checklist
- Find the top functions by invocations × duration × memory.
- Tune memory with Power Tuning, prefer
arm64where dependencies allow. - Filter and batch at the event source, avoid synchronous chains.
- Keep init code light, measure INIT duration since it is billed.
- Set log retention and log levels, check API Gateway and NAT costs.
- Use reserved concurrency, sane timeouts and budget alerts.
