Qwen 2.5 72B Instruct API Pricing

Qwen 2.5 72B Instruct by Alibaba costs $0.36 per million input tokens and $0.40 per million output tokens on DeepInfra.

Price verified from DeepInfra's pricing page.

Rates and limits

Input / 1M
$0.36
Output / 1M
$0.40
Cached input / 1M
Not listed
Batch discount
None
Context window
131K tokens
Max output
8K tokens
Knowledge cutoff
September 2024
Released
September 2024

Pricing note: Open-weight model; price is DeepInfra (Qwen2.5-72B-Instruct, served there with a 32k context). Alibaba Model Studio states the Qwen2.5 series is no longer available for calling, and Together AI removed Qwen2.5-72B-Instruct-Turbo from serverless on 2026-02-06.

What Qwen 2.5 72B Instruct costs per month

Token volume split 3:1 between input and output, with no prompt caching.

Tokens per monthInput / outputStandard
1M750K / 250K$0.37
10M7.5M / 2.5M$3.70
100M75M / 25M$37.00
Estimate your own workload with Qwen 2.5 72B Instruct

Where to run Qwen 2.5 72B Instruct

Prices above are DeepInfra's rates. Other platforms can charge differently; the cloud platform view compares them.

Benchmarks reported by Alibaba

MMLU
86.1%
HumanEval
86.6%
GPQA
49%

Qwen 2.5 72B Instruct pricing FAQ

How much does Qwen 2.5 72B Instruct cost per million tokens?

Qwen 2.5 72B Instruct costs $0.36 per million input tokens and $0.40 per million output tokens, as listed on DeepInfra's pricing page on September 30, 2026.

How much does Qwen 2.5 72B Instruct cost for 10 million tokens a month?

About $3.70 a month, assuming 7.5 million input and 2.5 million output tokens with no prompt caching.

What is the context window of Qwen 2.5 72B Instruct?

Qwen 2.5 72B Instruct accepts up to 131,072 tokens of context and returns up to 8,192 output tokens per request.

Does Qwen 2.5 72B Instruct support prompt caching?

No cached-input price is listed for Qwen 2.5 72B Instruct, so repeated prompt prefixes are billed at the standard input rate.

Is there a batch discount for Qwen 2.5 72B Instruct?

No batch API price is listed for Qwen 2.5 72B Instruct.

Where can I use the Qwen 2.5 72B Instruct API?

Qwen 2.5 72B Instruct is available through Fireworks AI and OpenRouter.

Prices shown as input / output per million tokens. All model prices · Side-by-side comparison