Llama 3.3 70B Instruct API Pricing
Llama 3.3 70B Instruct by Meta costs $1.04 per million input tokens and $1.04 per million output tokens on Together AI.
Price verified from Together AI's pricing page.
Rates and limits
- Input / 1M
- $1.04
- Output / 1M
- $1.04
- Cached input / 1M
- Not listed
- Batch discount
- 50% off
- Context window
- 128K tokens
- Max output
- 8K tokens
- Knowledge cutoff
- July 2024
- Released
- December 2024
Pricing note: Open-weight model; price is Together AI serverless (Llama-3.3-70B-Instruct-Turbo, FP8). Together lists no cached-input rate for it. Together Batch API lists it among the models at 50% off.
What Llama 3.3 70B Instruct costs per month
Token volume split 3:1 between input and output, with no prompt caching.
| Tokens per month | Input / output | Standard | Batch API |
|---|---|---|---|
| 1M | 750K / 250K | $1.04 | $0.52 |
| 10M | 7.5M / 2.5M | $10.40 | $5.20 |
| 100M | 75M / 25M | $104.00 | $52.00 |
Where to run Llama 3.3 70B Instruct
Prices above are Together AI's rates. Other platforms can charge differently; the cloud platform view compares them.
- AWS BedrockCloud / platform
- Google Cloud Vertex AICloud / platform
- Azure AI FoundryCloud / platform
- Together AICloud / platform
- Fireworks AICloud / platform
- ReplicateCloud / platform
- OpenRouterCloud / platform
- GroqCloud / platform
- Cloudflare Workers AICloud / platform
- Databricks Mosaic AICloud / platform
Benchmarks reported by Meta
- MMLU
- 86%
- HumanEval
- 88.4%
- GPQA
- 50.5%
- SWE-bench Verified
- 28.9%
Llama 3.3 70B Instruct pricing FAQ
How much does Llama 3.3 70B Instruct cost per million tokens?
Llama 3.3 70B Instruct costs $1.04 per million input tokens and $1.04 per million output tokens, as listed on Together AI's pricing page on September 30, 2026.
How much does Llama 3.3 70B Instruct cost for 10 million tokens a month?
About $10.40 a month, assuming 7.5 million input and 2.5 million output tokens with no prompt caching, or $5.20 through the batch API.
What is the context window of Llama 3.3 70B Instruct?
Llama 3.3 70B Instruct accepts up to 128,000 tokens of context and returns up to 8,192 output tokens per request.
Does Llama 3.3 70B Instruct support prompt caching?
No cached-input price is listed for Llama 3.3 70B Instruct, so repeated prompt prefixes are billed at the standard input rate.
Is there a batch discount for Llama 3.3 70B Instruct?
Yes. Requests sent through the batch API cost 50% less: $0.52 per million input tokens and $0.52 per million output tokens.
Where can I use the Llama 3.3 70B Instruct API?
Llama 3.3 70B Instruct is available through AWS Bedrock, Google Cloud Vertex AI, Azure AI Foundry, Together AI, Fireworks AI, Replicate, OpenRouter, Groq, Cloudflare Workers AI, and Databricks Mosaic AI.
Compare with other models
More from Meta
Prices shown as input / output per million tokens. All model prices · Side-by-side comparison