Polarison

Meta · open weights

Llama 3.3 70B Instruct API pricing by provider

Llama 3.3 70B Instruct is available from 10 providers. The cheapest is DeepInfra at $0.10 input / $0.32 output per 1M tokens (FP8). The priciest host, Together, charges 6.7x the cheapest.

Prices from the OpenRouter API, retrieved September 15, 2026. Weights: meta-llama/Llama-3.3-70B-Instruct.

ProviderPrecisionInput / 1MOutput / 1MDocument Q&A / monthContextMax output
DeepInfrafp8$0.10$0.32$49.60131K16K
Novitabf16$0.135$0.40$66.0012K11K
AkashMLfp8$0.20$0.52$95.60131K128K
Parasailfp8$0.22$0.50$103.00131K16K
SambaNova$0.45$0.90$207.00131K3K
Groq$0.59$0.79$259.70131K33K
CoreWeavefp16$0.71$0.71$305.30128K115K
Google$0.72$0.72$309.60128K8K
Cloudflarefp8$0.293$2.253$184.7924K22K
Together$1.04$1.04$447.20131K2K

Sorted by blended price (3 input : 1 output). Document Q&A assumes 50,000 questions a month with 8,000 input and 600 output tokens each. Precision is reported by the provider; “—” means not disclosed.

Frequently asked questions

What is the cheapest Llama 3.3 70B Instruct API provider?
DeepInfra is the cheapest healthy provider we track at $0.10 per 1M input tokens and $0.32 per 1M output tokens.
How much does Llama 3.3 70B Instruct cost?
Across 10 providers, input prices range from $0.10 to $1.04 per 1M tokens and output prices from $0.32 to $2.253.
What is Llama 3.3 70B Instruct’s context window?
Llama 3.3 70B Instruct supports up to 131K tokens, but some providers serve a smaller context window — check the table before choosing.

Other open-weight models