DeepSeek V4 Flash is DeepSeek's language model with a 1.3M context window and up to 1.0M output tokens, available from 16 providers, starting at $0.05 / 1M input and $0.16 / 1M output. An efficiency-optimized Mixture-of-Experts LLM from DeepSeek with 284B total and 13B activated parameters, supporting a 1M-token context window with reasoning and tool-use capabilities.
Specifications
Canonical IDdeepseek-v4-flash
TypeLanguage
StatusActive
CreatorDeepSeekDeepSeek
Providers
Context Window1.3M tokens
Max Output1.0M tokens
Input ModalitiesImagePDFText
Output ModalitiesText
Reasoning Effortsdefault
Parameters158B
HuggingFace Likes649
HuggingFace Downloads (30d)25,391
HuggingFace Downloads (all-time)25,391
Release Date · 4 months ago
Deprecation Date
Benchmarks
Intelligence Index
#31
Coding Index
#28
GPQA
#35
HLE
#38
IFBench
#201
Time to First Token
#112
SciCode
#43
LCR
#54
TerminalBench Hard
#77
TAU2
#34
Output TPS
#71

Capabilities

Input3/5
Text
Image
Audio·
Video·
PDF
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities7/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardBatch
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Alibaba Qwen logo
Alibaba Qwen
deepseek-v4-flash
$0.2$0.4$0.04$0.1$0.2N/A
Azure AI Foundry logo
Azure AI Foundry
azure_ai/deepseek-v4-flash
$0.19$0.51$0.028
DeepInfra logo
DeepInfra
deepinfra/deepseek-ai/DeepSeek-V4-Flash
$0.09$0.18$0.018
DeepSeek logo
DeepSeek
deepseek-v4-flash
$0.44$1.32$0.014$0.22$0.66$0.007
Fireworks AI logo
Fireworks AI
fireworks_ai/deepseek-v4-flash
$0.14$0.28$0.028
Hugging Face logo
Hugging Face
novita:deepseek/deepseek-v4-flash
$0.14$0.28N/A
Libertai
libertai/deepseek-v4-flash
$0.25$1.75N/A
Novita logo
Novita
novita/deepseek/deepseek-v4-flash
$0.14$0.28$0.028
OpenRouter logo
OpenRouter
~deepseek/deepseek-v4-flash-latest
$0.05$0.16$0.013
OpenRouter logo
OpenRouter
deepseek/deepseek-v4-flash
$0.0886$0.177$0.0177
Pinstripes
pinstripes/ps/deepseek-v4-flash
$0.1$0.2N/A
Qwen AI Platform
qwen_ai_platform/deepseek-v4-flash
$0.2$0.4$0.04
Qwencloud
qwencloud/deepseek-v4-flash
$0.2$0.4$0.04
Tencent logo
Tencent
tencent/deepseek-v4-flash
$0.14$0.28$0.0028
Tensormesh
tensormesh/deepseek-ai/DeepSeek-V4-Flash
$0.14$0.28N/A
Vercel AI Gateway logo
Vercel AI Gateway
deepseek/deepseek-v4-flash
$0.13$0.26$0.028
Weights & Biases logo
Weights & Biases
wandb/deepseek-ai/DeepSeek-V4-Flash
$0.14$0.28$0.07

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host DeepSeek V4 Flash, ranked by cheapest on-demand price. The model needs about 379 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g7e.24xlargeAWS4× RTX PRO Server 6000384 GB$16.57/hr
g4-standard-192GCP4× nvidia-rtx-pro-6000384 GB$18.00/hr
p4de.24xlargeAWS8× A100640 GB$27.45/hr
7 more instances can run DeepSeek V4 Flash
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
DeepSeek V4.731 Flash US$0.424$1.27Available
DeepSeek V4 420 FlashAvailable
DeepSeek V4 Flash Vision Exp1.0M$0.220$0.660Available
DeepSeek V4 Flash1.3M$0.050$0.160Current
DeepSeek V4 Flash VisionAvailable
DeepSeek V4 Flash 7311.0M$0.140$0.280Available
DeepSeek V4 Flash Thinking200K$0.250$1.75Available
DeepSeek V4 Flash US$0.200$0.400Available

Other Models

ModelTierReleasedContextInput / 1MOutput / 1M
DeepSeek V4 ProPro1.0M$0.660$1.98
DeepSeek V4 ProPro
DeepSeek V4 ProPro1.0M$0.435$0.870
DeepSeek V4 Pro 813Pro1.0M$1.32$3.96
DeepSeek V4 Pro USPro$2.40$4.80

Model IDs

~deepseek/deepseek-v4-flash-latest
accounts/fireworks/models/deepseek-v4-flash
azure_ai/deepseek-v4-flash
dashscope/deepseek-v4-flash
deepinfra/deepseek-ai/DeepSeek-V4-Flash
deepseek-ai/DeepSeek-V4-Flash
deepseek-llm-deepseek-v4-flash
deepseek-v4-flash
deepseek-v4-flash-high
deepseek-v4-flash-non-reasoning
deepseek-v4-flash(1)
deepseek-v4-flash*
deepseek/deepseek-v4-flash
deepseek/deepseek-v4-flash:free
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash
fireworks_ai/deepseek-v4-flash
libertai/deepseek-v4-flash
novita/deepseek/deepseek-v4-flash
pinstripes/ps/deepseek-v4-flash
qwen_ai_platform/deepseek-v4-flash
qwencloud/deepseek-v4-flash
tencent/deepseek-v4-flash
tensormesh/deepseek-ai/DeepSeek-V4-Flash
wandb/deepseek-ai/DeepSeek-V4-Flash