DeepSeek V4 731 Flash is DeepSeek's language model with a 1.3M context window and up to 944K output tokens, available from 17 providers, starting at $0.03 / 1M input and $0.153 / 1M output. An official release of DeepSeek's V4 Flash model with substantially enhanced agentic capabilities including reasoning, tool use, and implicit caching.
Specifications
Canonical IDdeepseek-v4-731-flash
TypeLanguage
StatusActive
CreatorDeepSeekDeepSeek
Providers
Context Window1.3M tokens
Max Output944K tokens
Input ModalitiesText
Output ModalitiesText
Reasoning Effortsdefault
Parameters304B
HuggingFace Likes1,499
HuggingFace Downloads (30d)15,366
HuggingFace Downloads (all-time)15,366
Release Date · 2 months ago
Deprecation Date

Capabilities

Input1/5
Text✓
Image·
Audio·
Video·
PDF·
Output1/5
Text✓
Image·
Audio·
Video·
Embedding·
Capabilities7/13
Reasoning✓
Adaptive Reasoning·
Function Calling✓
Parallel Function Calling✓
Structured Outputs✓
Native JSON Schema✓
Web Search✓
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching✓
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardFastBatch
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
Alibaba Qwen logo
Alibaba Qwen
deepseek-v4-flash-0731
$0.44$1.32$0.04———$0.22$0.66
Azure AI Foundry logo
Azure AI Foundry
azure_ai/DeepSeek-V4-Flash-0731
$0.44$1.32$0.014—————
Baseten
baseten/deepseek-ai/DeepSeek-V4-Flash-0731
$0.13$0.26$0.028—————
Cloudflare Workers AI logo
Cloudflare Workers AI
@cf/deepseek-ai/deepseek-v4-flash-0731
$0.44$1.32$0.014—————
DeepInfra logo
DeepInfra
deepinfra/deepseek-ai/DeepSeek-V4-Flash-0731
$0.08$0.18$0.016—————
Fireworks AI logo
Fireworks AI
fireworks_ai/deepseek-v4-flash-0731
$0.22$0.66$0.007$0.275$0.825$0.00875——
Hugging Face logo
Hugging Face
novita:deepseek/deepseek-v4-flash-0731
$0.44$1.32N/A—————
Hugging Face logo
Hugging Face
together_ai:deepseek-ai/DeepSeek-V4-Flash-0731
$0.14$0.28N/A—————
Nebius logo
Nebius
deepseek-ai/DeepSeek-V4-Flash-0731
$0.14$0.28N/A—————
Novita logo
Novita
novita/deepseek/deepseek-v4-flash-0731
$0.44$1.32$0.028—————
OpenRouter logo
OpenRouter
deepseek/deepseek-v4-flash-0731
$0.03$0.32$0.016—————
Perplexity logo
Perplexity
perplexity/perplexity/deepseek-v4-flash-0731
$0.13$0.26$0.028—————
Qwen AI Platform
qwen_ai_platform/deepseek-v4-flash-0731
$0.2$0.4$0.04—————
Qwencloud
qwencloud/deepseek-v4-flash-0731
$0.2$0.4$0.04—————
Scaleway logo
Scaleway
scaleway/deepseek-v4-flash-0731
$0.4$0.8$0.08—————
Together AI logo
Together AI
together_ai/deepseek-ai/DeepSeek-V4-Flash-0731
$0.14$0.28$0.03—————
Vercel AI Gateway logo
Vercel AI Gateway
deepseek/deepseek-v4-flash-0731
$0.076$0.153$0.014—————
Weights & Biases logo
Weights & Biases
wandb/deepseek-ai/DeepSeek-V4-Flash-0731
$0.13$0.28$0.07—————

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host DeepSeek V4 731 Flash, ranked by cheapest on-demand price. The model needs about 730 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
a3-ultragpu-8gGCP8× nvidia-h200-141gb1128 GB$13.82/hr
g7e.48xlargeAWS8× RTX PRO Server 6000768 GB$33.14/hr
g4-standard-384GCP8× nvidia-rtx-pro-6000768 GB$36.00/hr
7 more instances can run DeepSeek V4 731 Flash
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Model IDs

@cf/deepseek-ai/deepseek-v4-flash-0731
accounts/fireworks/models/deepseek-v4-flash-0731
azure_ai/deepseek-v4-flash-0731
azure_ai/DeepSeek-V4-Flash-0731
baseten/deepseek-ai/DeepSeek-V4-Flash-0731
dashscope/deepseek-v4-flash-0731
deepinfra/deepseek-ai/DeepSeek-V4-Flash-0731
deepseek-ai/DeepSeek-V4-Flash-0731
deepseek-v4-731-flash
deepseek-v4-flash-0731
deepseek/deepseek-v4-flash-0731
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731
fireworks_ai/deepseek-v4-flash-0731
nebius/deepseek-ai/DeepSeek-V4-Flash-0731
novita/deepseek/deepseek-v4-flash-0731
openrouter/deepseek/deepseek-v4-flash-0731
perplexity/perplexity/deepseek-v4-flash-0731
qwen_ai_platform/deepseek-v4-flash-0731
qwencloud/deepseek-v4-flash-0731
scaleway/deepseek-v4-flash-0731
together_ai/deepseek-ai/DeepSeek-V4-Flash-0731
wandb/deepseek-ai/DeepSeek-V4-Flash-0731