Gemma 4 Ultra 31B Instruct is Google's language model with a 131K context window, starting at $0.27 / 1M input and $0.76 / 1M output. An ultra-speed instruction-tuned variant of the Gemma 4 31B model, optimized for high-throughput inference.
Capabilities
Input2/5
Text✓
Image✓
Audio·
Video·
PDF·
Output1/5
Text✓
Image·
Audio·
Video·
Embedding·
Capabilities3/13
Reasoning✓
Adaptive Reasoning·
Function Calling✓
Parallel Function Calling·
Structured Outputs✓
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·
Pricing by Provider
US Dollar ($)
Per 1M tokens
| Provider | Standard | |
|---|---|---|
| Input $ / 1M | Output $ / 1M | |
| $0.27 | $0.76 | |
Cost Calculator
US Dollar ($)
Preset:
Cheapest Instances to Run It
Cloud GPU instances that can host Gemma 4 Ultra 31B Instruct, ranked by cheapest on-demand price. The model needs about 74 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.
All clouds
FP16 (full precision)
US Dollar ($)
Instance | Cloud | GPU | VRAM | Price | Cheapest region | |
|---|---|---|---|---|---|---|
| Standard_NP20s | 2× AMD Alveo U250 FPGA (64GB) | 128 GB | $3.30/hr | |||
| g7e.2xlarge | RTX PRO Server 6000 | 96 GB | $3.36/hr | |||
| Standard_NC24ads_A100_v4 | NVIDIA A100 | 80 GB | $3.67/hr | |||
Other Models
| Model | Tier | Released | Context | Input / 1M | Output / 1M |
|---|---|---|---|---|---|
| Gemma 4 31B | — | — | 256K | $0.140 | $0.400 |
| Gemma 4 26B A4B | — | — | 256K | $0.130 | $0.400 |
| Gemma 4 31B | — | — | — | — | — |
| Gemma 4 12B | — | — | — | — | — |
| Gemma 4 26B A4B | — | — | — | — | — |
| Gemma 4 E4B | — | — | — | — | — |
| Gemma 4 E2B | — | — | 128K | $0.040 | $0.080 |
| Gemma 4 | — | — | — | — | — |
| Gemma 4 E4B | — | — | — | — | — |
| Gemma 4 E2B | — | — | — | — | — |