Llama 3.1 405B Instruct API Pricing
Llama 3.1 405B Instruct by Meta costs $2.40 per million input tokens and $2.40 per million output tokens on AWS Bedrock.
Price verified from AWS Bedrock's pricing page.
Rates and limits
- Input / 1M
- $2.40
- Output / 1M
- $2.40
- Cached input / 1M
- Not listed
- Batch discount
- 50% off
- Context window
- 128K tokens
- Max output
- 4K tokens
- Knowledge cutoff
- December 2023
- Released
- July 2024
Pricing note: Open-weight model; price is AWS Bedrock on-demand in US West (Oregon), from the AWS Price List API ($0.0024 per 1K input/output tokens; batch $0.0012). Bedrock lists 405B only as batch in US East (N. Virginia). Together AI removed it from serverless on 2026-02-06.
What Llama 3.1 405B Instruct costs per month
Token volume split 3:1 between input and output, with no prompt caching.
| Tokens per month | Input / output | Standard | Batch API |
|---|---|---|---|
| 1M | 750K / 250K | $2.40 | $1.20 |
| 10M | 7.5M / 2.5M | $24.00 | $12.00 |
| 100M | 75M / 25M | $240.00 | $120.00 |
Where to run Llama 3.1 405B Instruct
Prices above are AWS Bedrock's rates. Other platforms can charge differently; the cloud platform view compares them.
- AWS BedrockCloud / platform
- Google Cloud Vertex AICloud / platform
- Azure AI FoundryCloud / platform
- Fireworks AICloud / platform
- ReplicateCloud / platform
- OpenRouterCloud / platform
- Databricks Mosaic AICloud / platform
Benchmarks reported by Meta
- MMLU
- 88.6%
- HumanEval
- 89%
- GPQA
- 51.1%
- HellaSwag
- 95.4%
Llama 3.1 405B Instruct pricing FAQ
How much does Llama 3.1 405B Instruct cost per million tokens?
Llama 3.1 405B Instruct costs $2.40 per million input tokens and $2.40 per million output tokens, as listed on AWS Bedrock's pricing page on September 30, 2026.
How much does Llama 3.1 405B Instruct cost for 10 million tokens a month?
About $24.00 a month, assuming 7.5 million input and 2.5 million output tokens with no prompt caching, or $12.00 through the batch API.
What is the context window of Llama 3.1 405B Instruct?
Llama 3.1 405B Instruct accepts up to 128,000 tokens of context and returns up to 4,096 output tokens per request.
Does Llama 3.1 405B Instruct support prompt caching?
No cached-input price is listed for Llama 3.1 405B Instruct, so repeated prompt prefixes are billed at the standard input rate.
Is there a batch discount for Llama 3.1 405B Instruct?
Yes. Requests sent through the batch API cost 50% less: $1.20 per million input tokens and $1.20 per million output tokens.
Where can I use the Llama 3.1 405B Instruct API?
Llama 3.1 405B Instruct is available through AWS Bedrock, Google Cloud Vertex AI, Azure AI Foundry, Fireworks AI, Replicate, OpenRouter, and Databricks Mosaic AI.
Compare with other models
More from Meta
Prices shown as input / output per million tokens. All model prices · Side-by-side comparison