Gemma 4 31B pricing and routing profile
Gemma 4 31B from Google lists at $0.09 per million input tokens and $0.34 per million output tokens, with a 262,144-token context window. Prices from OpenRouter's catalogue on 2026-09-28.
Tool callingStructured outputReasoning controlLog probabilities
Cost per 1,000 requests
| Request shape | List cost |
|---|---|
| Classification or routing (600 tokens in, 10 out) | $0.0574 |
| Chat reply (2,000 in, 400 out) | $0.316 |
| Retrieval-augmented answer (8,000 in, 500 out) | $0.89 |
Before prompt caching. Try your own shape in the cost calculator.
Can Gimbix route to it first?
Yes. 5 of 15 tracked hosts return token log probabilities for Gemma 4 31B, which is the confidence signal Gimbix uses to let a cheaper model answer first.
Measured by Gimbix
On 400 real BANKING77 customer messages, Gemma 4 31B answered 304 correctly (95% interval 71.6%–79.9%), with 1 invalid answers, at $0.0381 per 1,000 requests and a median of 791 ms. See all 22 models.
Who serves Gemma 4 31B
| Host | Input / output per 1M | Log probabilities | Uptime, last 30 min |
|---|---|---|---|
| Reka | $0.08 / $0.30 | yes | 100.0% |
| DeepInfra | $0.09 / $0.34 | no | 99.3% |
| CoreWeave | $0.10 / $0.34 | yes | 97.0% |
| Venice | $0.12 / $0.36 | yes | 99.2% |
| Chutes | $0.12 / $0.37 | no | 98.8% |
| DeepInfra | $0.13 / $0.38 | no | 96.0% |
| Crusoe | $0.14 / $0.40 | no | 100.0% |
| Friendli | $0.14 / $0.40 | no | 99.9% |
| Novita | $0.14 / $0.40 | yes | 88.9% |
| Parasail | $0.15 / $0.40 | yes | 99.5% |
| DeepInfra | $0.27 / $0.76 | no | 93.8% |
| Io Net | $0.38 / $1.15 | yes | 98.5% |
| SambaNova | $0.38 / $1.15 | no | 90.9% |
| ModelRun | $0.75 / $1.00 | no | 99.9% |
| SiliconFlow | $0.75 / $1.00 | no | 99.7% |
A host that lists log probabilities may still omit them; Gimbix checks each one before relying on it.
Cheaper models that can answer first
Models priced below Gemma 4 31B that expose a confidence signal. Whether they answer your traffic as well is what a Gimbix pilot measures.
| Model | Input / output per 1M | Cheaper on a chat reply |
|---|---|---|
| Qwen3 Coder 30B A3B Instruct | $0.07 / $0.28 | 20% |
| Nemotron 3.5 Lightning | $0.08 / $0.20 | 24% |
| Llama 3.2 3B Instruct | $0.05 / $0.33 | 27% |
| Gemma 4 26B A4B | $0.0675 / $0.225 | 29% |
| Granite 4.2 8B | $0.06 / $0.25 | 30% |
| Nemotron 3 Nano 30B A3B | $0.05 / $0.20 | 43% |
Compare Gemma 4 31B
- Gemma 4 31B vs Claude Opus 5.5
- Gemma 4 31B vs Claude Sonnet 5
- Gemma 4 31B vs Claude Fable 5.1
- Gemma 4 31B vs Claude Opus 5
- Gemma 4 31B vs GPT-6 Sol
- Gemma 4 31B vs GPT-6 Astra
- Gemma 4 31B vs GPT-5.6 Sol
- Gemma 4 31B vs GPT-5.6 Terra
- Gemma 4 31B vs GPT-5.5
- Gemma 4 31B vs GPT-5.4
- Gemma 4 31B vs Gemini 3.1 Pro Preview
- Gemma 4 31B vs Gemini 3.8 Flash
- Gemma 4 31B vs Grok 4.7
- Gemma 4 31B vs Qwen3.8 Max (0902)
- Gemma 4 31B vs Kimi K3
- Gemma 4 31B vs GLM 5.3
- Gemma 4 31B vs DeepSeek V4 Pro 0423
- Gemma 4 31B vs Mistral Large 3 2512
- Gemma 4 31B vs Nova Premier 1.0
- Gemma 4 31B vs Command A
- Gemma 4 31B vs GPT-6 Luna
- Gemma 4 31B vs GPT-5.6 Luna
- Gemma 4 31B vs GPT-5.4 Mini
- Gemma 4 31B vs GPT-5.4 Nano
- Gemma 4 31B vs GPT-4.1 Mini
- Gemma 4 31B vs gpt-oss-120b
- Gemma 4 31B vs gpt-oss-20b
- Gemma 4 31B vs Claude Haiku 4.5
- Gemma 4 31B vs Gemini 3.5 Flash Lite
- Gemma 4 31B vs DeepSeek V4.1 Flash
- Gemma 4 31B vs DeepSeek V4 Flash 0731
- Gemma 4 31B vs GLM 5.3 Flash
- Gemma 4 31B vs Hy4 preview
- Gemma 4 31B vs MiMo-V2.6-Flash
- Gemma 4 31B vs Nemotron 3.5 Lightning
- Gemma 4 31B vs Nemotron 3 Ultra
- Gemma 4 31B vs Qwen3.8 Flash
- Gemma 4 31B vs Qwen3.7 Flash
- Gemma 4 31B vs Qwen3.8 27B
- Gemma 4 31B vs MiniMax M3
- Gemma 4 31B vs Kimi K2.6
- Gemma 4 31B vs Mistral Small 4
- Gemma 4 31B vs Ministral 3 8B 2512
- Gemma 4 31B vs Llama 4 Maverick
- Gemma 4 31B vs Llama 3.3 70B Instruct
- Gemma 4 31B vs Nova Lite 1.0
- Gemma 4 31B vs Nova Micro 1.0
- Gemma 4 31B vs Command R7B (12-2024)
- Gemma 4 31B vs Phi 4
Questions
How much does Gemma 4 31B cost?
Gemma 4 31B lists at $0.09 per million input tokens and $0.34 per million output tokens on OpenRouter as of 2026-09-28. A typical chat reply of 2,000 input and 400 output tokens costs about $0.316 per 1,000 requests.
What is the context window of Gemma 4 31B?
262,144 tokens.
Can a cheaper model replace Gemma 4 31B?
Sometimes. Gimbix proves it on your own traffic before routing: the cheaper model must match Gemma 4 31B on one half of your requests and make at most three answers worse on the other.