DeepSeek V4 731 Flash is DeepSeek's language model with a 1.3M context window and up to 944K output tokens, available from 12 providers, starting at $0.065 / 1M input and $0.153 / 1M output. An official release of DeepSeek's V4 Flash model with substantially enhanced agentic capabilities including reasoning, tool use, and implicit caching.
Specifications
Canonical IDdeepseek-v4-731-flash
TypeLanguage
StatusActive
CreatorDeepSeekDeepSeek
Providers
Context Window1.3M tokens
Max Output944K tokens
Input ModalitiesText
Output ModalitiesText
Reasoning Effortsdefault
Parameters304B
HuggingFace Likes1,499
HuggingFace Downloads (30d)15,366
HuggingFace Downloads (all-time)15,366
Release Date · 1 month ago

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities7/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema
Web Search
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardBatch
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Alibaba Qwen logo
Alibaba Qwen
deepseek-v4-flash-0731
$0.44$1.32$0.04$0.22$0.66N/A
Cloudflare Workers AI logo
Cloudflare Workers AI
@cf/deepseek-ai/deepseek-v4-flash-0731
$0.44$1.32$0.014
DeepInfra logo
DeepInfra
deepinfra/deepseek-ai/DeepSeek-V4-Flash-0731
$0.08$0.18$0.016
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731
$0.22$0.66$0.007
Fireworks AI logo
Fireworks AI
fireworks_ai/deepseek-v4-flash-0731
$0.14$0.28$0.028
Hugging Face logo
Hugging Face
novita:deepseek/deepseek-v4-flash-0731
$0.44$1.32N/A
Hugging Face logo
Hugging Face
together_ai:deepseek-ai/DeepSeek-V4-Flash-0731
$0.14$0.28N/A
Nebius logo
Nebius
deepseek-ai/DeepSeek-V4-Flash-0731
$0.14$0.28N/A
Novita logo
Novita
novita/deepseek/deepseek-v4-flash-0731
$0.44$1.32$0.028
OpenRouter logo
OpenRouter
deepseek/deepseek-v4-flash-0731
$0.065$0.18$0.016N/AN/AN/A
OpenRouter logo
OpenRouter
deepseek/deepseek-v4-flash-0731:batch
$0.14$0.28$0.03
Perplexity logo
Perplexity
perplexity/perplexity/deepseek-v4-flash-0731
$0.13$0.26$0.028
Together AI logo
Together AI
together_ai/deepseek-ai/DeepSeek-V4-Flash-0731
$0.14$0.28$0.03
Vercel AI Gateway logo
Vercel AI Gateway
deepseek/deepseek-v4-flash-0731
$0.076$0.153$0.014
Weights & Biases logo
Weights & Biases
wandb/deepseek-ai/DeepSeek-V4-Flash-0731
$0.13$0.28$0.07

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host DeepSeek V4 731 Flash, ranked by cheapest on-demand price. The model needs about 730 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g7e.48xlargeAWS8× RTX PRO Server 6000768 GB$33.14/hr
g4-standard-384GCP8× nvidia-rtx-pro-6000768 GB$36.00/hr
Standard_ND96isr_MI300X_v5Azure8× AMD Instinct MI300X1536 GB$48.00/hr
7 more instances can run DeepSeek V4 731 Flash
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Model IDs

@cf/deepseek-ai/deepseek-v4-flash-0731
accounts/fireworks/models/deepseek-v4-flash-0731
dashscope/deepseek-v4-flash-0731
deepinfra/deepseek-ai/DeepSeek-V4-Flash-0731
deepseek-ai/DeepSeek-V4-Flash-0731
deepseek-v4-731-flash
deepseek-v4-flash-0731
deepseek/deepseek-v4-flash-0731
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731
fireworks_ai/deepseek-v4-flash-0731
novita/deepseek/deepseek-v4-flash-0731
perplexity/perplexity/deepseek-v4-flash-0731
together_ai/deepseek-ai/DeepSeek-V4-Flash-0731
wandb/deepseek-ai/DeepSeek-V4-Flash-0731