Gemma 2 9B Instruct is Google's language model with a 8K context window and up to 8K output tokens, available from 2 providers, starting at $0.2 / 1M input and $0.2 / 1M output. An instruction-tuned 9B Gemma 2 LLM with strong performance on reasoning and language tasks.
Specifications
Canonical IDgoogle-gemma-2-9b-instruct
TypeLanguage
StatusActive
CreatorGoogleGoogle
Providers
Context Window8K tokens
Max Output8K tokens
Input ModalitiesImageText
Output ModalitiesText
Parameters9B

Capabilities

Input2/5
Text✓
Image✓
Audio·
Video·
PDF·
Output1/5
Text✓
Image·
Audio·
Video·
Embedding·
Capabilities1/13
Reasoning·
Adaptive Reasoning·
Function Calling✓
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/gemma2-9b-it
$0.2$0.2
Google Gemini logo
Google Gemini
gemini/gemini-gemma-2-9b-it
$0.35$1.05

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Gemma 2 9B Instruct, ranked by cheapest on-demand price. The model needs about 22 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g2-standard-4GCPnvidia-l424 GB$0.705/hr
g6.xlargeAWSL422 GB$0.805/hr
g2-standard-8GCPnvidia-l424 GB$0.851/hr
7 more instances can run Gemma 2 9B Instruct
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Model IDs

accounts/fireworks/models/gemma2-9b-it
fireworks_ai/accounts/fireworks/models/gemma2-9b-it
gemini/gemini-gemma-2-9b-it
google-gemma-2-9b-instruct
huggingface-llm-gemma-2-9b-instruct