Polarison

Google · open weights

Gemma 4 31B API pricing by provider

Gemma 4 31B is available from 11 providers. The cheapest is DeepInfra at $0.09 input / $0.34 output per 1M tokens (FP4). The priciest host, SiliconFlow, charges 5.3x the cheapest.

Prices from the OpenRouter API, retrieved September 15, 2026. Weights: google/gemma-4-31B-it.

ProviderPrecisionInput / 1MOutput / 1MDocument Q&A / monthContextMax output
DeepInfrafp4$0.09$0.34$46.20262K16K
CoreWeavefp4$0.10$0.34$50.20262K236K
Venicebf16$0.12$0.36$58.80256K8K
Chutesfp4$0.12$0.37$59.10131K66K
DeepInfrafp8$0.13$0.38$63.40262K16K
Crusoe$0.14$0.40$68.00262K262K
Friendli$0.14$0.40$68.00262K8K
Parasailfp8$0.15$0.40$72.00262K236K
DeepInfrafp8$0.27$0.76$130.80131K8K
Together$0.39$0.97$185.10262K236K
SambaNova$0.38$1.15$186.50131K118K
ModelRunfp4$0.75$1.00$330.00262K236K
SiliconFlowfp8$0.75$1.00$330.00262K236K

Sorted by blended price (3 input : 1 output). Document Q&A assumes 50,000 questions a month with 8,000 input and 600 output tokens each. Precision is reported by the provider; “—” means not disclosed.

Frequently asked questions

What is the cheapest Gemma 4 31B API provider?
DeepInfra is the cheapest healthy provider we track at $0.09 per 1M input tokens and $0.34 per 1M output tokens.
How much does Gemma 4 31B cost?
Across 11 providers, input prices range from $0.09 to $0.75 per 1M tokens and output prices from $0.34 to $1.15.
What is Gemma 4 31B’s context window?
Gemma 4 31B supports up to 262K tokens, but some providers serve a smaller context window — check the table before choosing.

Other open-weight models