Gemma 4 Ultra 31B Instruct is Google's language model with a 131K context window, starting at $0.27 / 1M input and $0.76 / 1M output. An ultra-speed instruction-tuned variant of the Gemma 4 31B model, optimized for high-throughput inference.
Specifications
Canonical IDgoogle-gemma-4-ultra-31b-instruct
TypeLanguage
StatusActive
CreatorGoogleGoogle
Providers
Context Window131K tokens
Input ModalitiesImageText
Output ModalitiesText
Reasoning Effortsdefault
Parameters31B

Capabilities

Input2/5
Text
Image
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities3/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling·
Structured Outputs
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
DeepInfra logo
DeepInfra
deepinfra/google/gemma-4-31B-it-Ultra
$0.27$0.76

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Gemma 4 Ultra 31B Instruct, ranked by cheapest on-demand price. The model needs about 74 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_NP20sAzure2× AMD Alveo U250 FPGA (64GB)128 GB$3.30/hr
g7e.2xlargeAWSRTX PRO Server 600096 GB$3.36/hr
Standard_NC24ads_A100_v4AzureNVIDIA A10080 GB$3.67/hr
7 more instances can run Gemma 4 Ultra 31B Instruct
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Other Models

ModelTierReleasedContextInput / 1MOutput / 1M
Gemma 4 31B256K$0.140$0.400
Gemma 4 26B A4B256K$0.130$0.400
Gemma 4 31B
Gemma 4 12B
Gemma 4 26B A4B
Gemma 4 E4B
Gemma 4 E2B128K$0.040$0.080
Gemma 4
Gemma 4 E4B
Gemma 4 E2B

Model IDs

deepinfra/google/gemma-4-31B-it-Ultra
google-gemma-4-ultra-31b-instruct
google/gemma-4-31B-it-Ultra