DeepSeek · open weights
DeepSeek V4.1 Flash API pricing by provider
DeepSeek V4.1 Flash is available from 16 providers. The cheapest is Relace at $0.15 input / $0.60 output per 1M tokens (FP4). DeepSeek’s own API charges $0.30 / $1.20, 50% more than the cheapest host. The priciest host, Venice, charges 2.5x the cheapest.
Prices from the OpenRouter API, retrieved September 15, 2026. Weights: deepseek-ai/DeepSeek-V4.1-Flash.
| Provider | Precision | Input / 1M | Output / 1M | Document Q&A / month | Context | Max output |
|---|---|---|---|---|---|---|
| Relace | fp4 | $0.15 | $0.60 | $78.00 | 1.05M | 944K |
| DeepInfra | fp8 | $0.20 | $0.60 | $98.00 | 1.05M | 131K |
| Morph | fp8 | $0.225 | $0.90 | $117.00 | 1.05M | 944K |
| Reka | fp4 | $0.29 | $1.16 | $150.80 | 262K | 131K |
| Alibaba | — | $0.30 | $1.20 | $156.00 | 1M | 393K |
| Together | — | $0.30 | $1.20 | $156.00 | 1.05M | 944K |
| SiliconFlow | fp8 | $0.30 | $1.20 | $156.00 | 1.05M | 393K |
| Modal | — | $0.30 | $1.20 | $156.00 | 1.05M | 944K |
| Wafer | — | $0.30 | $1.20 | $156.00 | 1.05M | 944K |
| BaseTen | fp8 | $0.30 | $1.20 | $156.00 | 1.05M | 33K |
| Parasail | fp8 | $0.30 | $1.20 | $156.00 | 1.05M | 944K |
| GMICloud | fp8 | $0.30 | $1.20 | $156.00 | 1.05M | 944K |
| Novita | fp8 | $0.30 | $1.20 | $156.00 | 1.05M | 393K |
| DeepSeekcreator | — | $0.30 | $1.20 | $156.00 | 1.05M | 384K |
| Phala | — | $0.345 | $1.38 | $179.40 | 1.05M | 393K |
| Venice | fp8 | $0.375 | $1.50 | $195.00 | 1M | 131K |
Sorted by blended price (3 input : 1 output). Document Q&A assumes 50,000 questions a month with 8,000 input and 600 output tokens each. Precision is reported by the provider; “—” means not disclosed.
Compare with DeepSeek’s API page
See benchmarks, cached-input pricing and comparisons for DeepSeek V4.1 Flash on its model page.
Frequently asked questions
- What is the cheapest DeepSeek V4.1 Flash API provider?
- Relace is the cheapest healthy provider we track at $0.15 per 1M input tokens and $0.60 per 1M output tokens.
- How much does DeepSeek V4.1 Flash cost?
- Across 16 providers, input prices range from $0.15 to $0.375 per 1M tokens and output prices from $0.60 to $1.50.
- What is DeepSeek V4.1 Flash’s context window?
- DeepSeek V4.1 Flash supports up to 1.05M tokens, but some providers serve a smaller context window — check the table before choosing.
Other open-weight models
- DeepSeek V4 Pro 0813from $0.96 / $2.88
- Kimi K3from $2.10 / $10.95
- GLM 5.3from $0.91 / $2.86
- GLM 5.3 Flashfrom $0.075 / $0.25
- Qwen3.8 27Bfrom $0.15 / $2.00
- MiniMax M3from $0.23 / $0.96
- gpt-oss-120bfrom $0.03 / $0.17
- gpt-oss-20bfrom $0.02 / $0.10
- Llama 4 Maverickfrom $0.1875 / $0.6525
- Llama 3.3 70B Instructfrom $0.10 / $0.32
- Gemma 4 31Bfrom $0.09 / $0.34
- Mistral Small 4from $0.15 / $0.60
- Nemotron 3 Superfrom $0.085 / $0.40