GPT-4.1 API Pricing
GPT-4.1 by OpenAI costs $2.00 per million input tokens and $8.00 per million output tokens.
Price verified from OpenAI's pricing page.
Rates and limits
- Input / 1M
- $2.00
- Output / 1M
- $8.00
- Cached input / 1M
- $0.50
- Batch discount
- 50% off
- Context window
- 1.05M tokens
- Max output
- 32K tokens
- Knowledge cutoff
- October 2024
- Released
- April 2025
What GPT-4.1 costs per month
Token volume split 3:1 between input and output, with no prompt caching.
| Tokens per month | Input / output | Standard | Batch API |
|---|---|---|---|
| 1M | 750K / 250K | $3.50 | $1.75 |
| 10M | 7.5M / 2.5M | $35.00 | $17.50 |
| 100M | 75M / 25M | $350.00 | $175.00 |
Where to run GPT-4.1
Prices above are OpenAI's direct rates. Cloud platforms can charge differently; the cloud platform view compares them.
- OpenAI APIDirect API
- Azure OpenAI ServiceCloud / platform
- Azure AI FoundryCloud / platform
- OpenRouterCloud / platform
Benchmarks reported by OpenAI
- MMLU
- 89%
- HumanEval
- 92%
- GPQA
- 60%
- MMMU
- 75%
GPT-4.1 pricing FAQ
How much does GPT-4.1 cost per million tokens?
GPT-4.1 costs $2.00 per million input tokens and $8.00 per million output tokens, as listed on OpenAI's pricing page on September 30, 2026.
How much does GPT-4.1 cost for 10 million tokens a month?
About $35.00 a month, assuming 7.5 million input and 2.5 million output tokens with no prompt caching, or $17.50 through the batch API.
What is the context window of GPT-4.1?
GPT-4.1 accepts up to 1,047,576 tokens of context and returns up to 32,000 output tokens per request.
Does GPT-4.1 support prompt caching?
Yes. Cached input tokens cost $0.50 per million, 75% less than the standard input rate.
Is there a batch discount for GPT-4.1?
Yes. Requests sent through the batch API cost 50% less: $1.00 per million input tokens and $4.00 per million output tokens.
Where can I use the GPT-4.1 API?
GPT-4.1 is available through OpenAI API, Azure OpenAI Service, Azure AI Foundry, and OpenRouter.
Compare with other models
More from OpenAI
Prices shown as input / output per million tokens. All model prices · Side-by-side comparison