Gemma 3N E4B IT is Google's language model with a 33K context window and up to 2K output tokens, starting at $0.06 / 1M input and $0.12 / 1M output. A multimodal instruction-tuned Gemma 3N model at an effective 4B parameter size, designed for mobile and low-resource devices supporting text, visual, and audio inputs.
Specifications
Canonical IDgoogle-gemma-3n-e4b-it
TypeLanguage
StatusDeprecated
CreatorGoogleGoogle
Providers
Context Window33K tokens
Max Output2K tokens
Input ModalitiesText
Output ModalitiesText
Parameters7.85B
HuggingFace Likes906
HuggingFace Downloads (30d)44,286
HuggingFace Downloads (all-time)1,257,738
Release Date · 1 year ago
Knowledge Cutoff · 2 years ago
Deprecation Date

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities2/13
Reasoning·
Adaptive Reasoning·
Function Calling·
Parallel Function Calling·
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Together AI logo
Together AI
together_ai/google/gemma-3n-E4B-it
$0.06$0.12

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Gemma 3N E4B IT, ranked by cheapest on-demand price. The model needs about 19 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g2-standard-4GCPnvidia-l424 GB$0.705/hr
g6.xlargeAWSL422 GB$0.805/hr
g2-standard-8GCPnvidia-l424 GB$0.851/hr
7 more instances can run Gemma 3N E4B IT
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Model IDs

google-gemma-3n-e4b-it
google/gemma-3n-e4b-it
google/gemma-3n-e4b-it:free
together_ai/google/gemma-3n-E4B-it