Polarison

DeepSeek · open weights

DeepSeek V4.1 Flash API pricing by provider

DeepSeek V4.1 Flash is available from 16 providers. The cheapest is Relace at $0.15 input / $0.60 output per 1M tokens (FP4). DeepSeek’s own API charges $0.30 / $1.20, 50% more than the cheapest host. The priciest host, Venice, charges 2.5x the cheapest.

Prices from the OpenRouter API, retrieved September 15, 2026. Weights: deepseek-ai/DeepSeek-V4.1-Flash.

ProviderPrecisionInput / 1MOutput / 1MDocument Q&A / monthContextMax output
Relacefp4$0.15$0.60$78.001.05M944K
DeepInfrafp8$0.20$0.60$98.001.05M131K
Morphfp8$0.225$0.90$117.001.05M944K
Rekafp4$0.29$1.16$150.80262K131K
Alibaba$0.30$1.20$156.001M393K
Together$0.30$1.20$156.001.05M944K
SiliconFlowfp8$0.30$1.20$156.001.05M393K
Modal$0.30$1.20$156.001.05M944K
Wafer$0.30$1.20$156.001.05M944K
BaseTenfp8$0.30$1.20$156.001.05M33K
Parasailfp8$0.30$1.20$156.001.05M944K
GMICloudfp8$0.30$1.20$156.001.05M944K
Novitafp8$0.30$1.20$156.001.05M393K
DeepSeekcreator$0.30$1.20$156.001.05M384K
Phala$0.345$1.38$179.401.05M393K
Venicefp8$0.375$1.50$195.001M131K

Sorted by blended price (3 input : 1 output). Document Q&A assumes 50,000 questions a month with 8,000 input and 600 output tokens each. Precision is reported by the provider; “—” means not disclosed.

Compare with DeepSeek’s API page

See benchmarks, cached-input pricing and comparisons for DeepSeek V4.1 Flash on its model page.

Frequently asked questions

What is the cheapest DeepSeek V4.1 Flash API provider?
Relace is the cheapest healthy provider we track at $0.15 per 1M input tokens and $0.60 per 1M output tokens.
How much does DeepSeek V4.1 Flash cost?
Across 16 providers, input prices range from $0.15 to $0.375 per 1M tokens and output prices from $0.60 to $1.50.
What is DeepSeek V4.1 Flash’s context window?
DeepSeek V4.1 Flash supports up to 1.05M tokens, but some providers serve a smaller context window — check the table before choosing.

Other open-weight models