Llama 3.1 405B Instruct API Pricing

Llama 3.1 405B Instruct by Meta costs $2.40 per million input tokens and $2.40 per million output tokens on AWS Bedrock.

Price verified from AWS Bedrock's pricing page.

Rates and limits

Input / 1M
$2.40
Output / 1M
$2.40
Cached input / 1M
Not listed
Batch discount
50% off
Context window
128K tokens
Max output
4K tokens
Knowledge cutoff
December 2023
Released
July 2024

Pricing note: Open-weight model; price is AWS Bedrock on-demand in US West (Oregon), from the AWS Price List API ($0.0024 per 1K input/output tokens; batch $0.0012). Bedrock lists 405B only as batch in US East (N. Virginia). Together AI removed it from serverless on 2026-02-06.

What Llama 3.1 405B Instruct costs per month

Token volume split 3:1 between input and output, with no prompt caching.

Tokens per monthInput / outputStandardBatch API
1M750K / 250K$2.40$1.20
10M7.5M / 2.5M$24.00$12.00
100M75M / 25M$240.00$120.00
Estimate your own workload with Llama 3.1 405B Instruct

Where to run Llama 3.1 405B Instruct

Prices above are AWS Bedrock's rates. Other platforms can charge differently; the cloud platform view compares them.

Benchmarks reported by Meta

MMLU
88.6%
HumanEval
89%
GPQA
51.1%
HellaSwag
95.4%

Llama 3.1 405B Instruct pricing FAQ

How much does Llama 3.1 405B Instruct cost per million tokens?

Llama 3.1 405B Instruct costs $2.40 per million input tokens and $2.40 per million output tokens, as listed on AWS Bedrock's pricing page on September 30, 2026.

How much does Llama 3.1 405B Instruct cost for 10 million tokens a month?

About $24.00 a month, assuming 7.5 million input and 2.5 million output tokens with no prompt caching, or $12.00 through the batch API.

What is the context window of Llama 3.1 405B Instruct?

Llama 3.1 405B Instruct accepts up to 128,000 tokens of context and returns up to 4,096 output tokens per request.

Does Llama 3.1 405B Instruct support prompt caching?

No cached-input price is listed for Llama 3.1 405B Instruct, so repeated prompt prefixes are billed at the standard input rate.

Is there a batch discount for Llama 3.1 405B Instruct?

Yes. Requests sent through the batch API cost 50% less: $1.20 per million input tokens and $1.20 per million output tokens.

Where can I use the Llama 3.1 405B Instruct API?

Llama 3.1 405B Instruct is available through AWS Bedrock, Google Cloud Vertex AI, Azure AI Foundry, Fireworks AI, Replicate, OpenRouter, and Databricks Mosaic AI.

Prices shown as input / output per million tokens. All model prices · Side-by-side comparison