DeepSeek-OCR 2 is DeepSeek's image to text model with a 8K context window and up to 8K output tokens, starting at $0.03 / 1M input and $0.03 / 1M output. An upgraded multimodal document recognition model from DeepSeek AI featuring the DeepEncoder V2 architecture for improved document understanding.
Specifications
Canonical IDdeepseek-ocr-2
TypeImage to Text
StatusActive
CreatorDeepSeekDeepSeek
Providers
Context Window8K tokens
Max Output8K tokens
Input ModalitiesImageText
Output ModalitiesText

Capabilities

Input2/5
Text
Image
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities0/13
Reasoning·
Adaptive Reasoning·
Function Calling·
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Novita logo
Novita
novita/deepseek/deepseek-ocr-2
$0.03$0.03

Cost Calculator

US Dollar ($)
Preset:

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Mistral OCR 4Available
Mistral OCR 4 AnnotAvailable
DeepSeek-OCR 28K$0.030$0.030Current
LightOn OCR 2 1BAvailable
Qianfan OCR Fast66KDeprecated
DeepSeek-OCRAvailable
Document OCRAvailable
Mistral OCRDeprecated
OCRAvailable
Prebuilt DocumentAvailable
Prebuilt LayoutAvailable

Model IDs

deepseek-ocr-2
deepseek/deepseek-ocr-2
novita/deepseek/deepseek-ocr-2