Gemma 4 Ultra 31B Instruct is Google's language model. An ultra-speed instruction-tuned variant of the Gemma 4 31B model, optimized for high-throughput inference.
Specifications
Canonical IDgoogle-gemma-4-ultra-31b-instruct
TypeLanguage
StatusActive
CreatorGoogleGoogle
Input ModalitiesText
Output ModalitiesText
Parameters31B

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities0/13
Reasoning·
Adaptive Reasoning·
Function Calling·
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Cheapest Instances to Run It

Cloud GPU instances that can host Gemma 4 Ultra 31B Instruct, ranked by cheapest on-demand price. The model needs about 74 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_NP20sAzure2× AMD Alveo U250 FPGA (64GB)128 GB$3.30/hr
g7e.2xlargeAWSRTX PRO Server 600096 GB$3.36/hr
g2-standard-48GCP4× nvidia-l496 GB$3.99/hr
7 more instances can run Gemma 4 Ultra 31B Instruct
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Other Models

ModelTierReleasedContextInput / 1MOutput / 1M
Gemma 4 31B256K$0.140$0.400
Gemma 4 26B A4B256K$0.130$0.400
Gemma 4 31B
Gemma 4 12B
Gemma 4 26B A4B
Gemma 4 E4B
Gemma 4 E2B128K$0.040$0.080
Gemma 4
Gemma 4 E4B
Gemma 4 E2B

Model IDs

google-gemma-4-ultra-31b-instruct
google/gemma-4-31B-it-Ultra