Step 3 VL 10B is StepFun's language model. A 10B-parameter vision-language model from StepFun capable of understanding and reasoning over both images and text.
Specifications
Canonical IDstep-3-10b-vl
TypeLanguage
StatusActive
CreatorStepFunStepFun
Input ModalitiesText
Output ModalitiesText
Parameters10B
Benchmarks
Intelligence Index
9.3
#319
GPQA
0.7
#225
HLE
0.1
#170
IFBench
0.5
#156
Time to First Token
0.00s
#341
SciCode
0.3
#242
LCR
0.0
#422
TerminalBench Hard
0.1
#256
TAU2
0.2
#335
Output TPS
0.0
#463

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities0/13
Reasoning·
Adaptive Reasoning·
Function Calling·
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Cheapest Instances to Run It

Cloud GPU instances that can host Step 3 VL 10B, ranked by cheapest on-demand price. The model needs about 24 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g2-standard-4GCPnvidia-l424 GB$0.705/hrus-east4
g2-standard-8GCPnvidia-l424 GB$0.851/hrus-east4
Standard_NV16as_v4AzureAMD Radeon Instinct MI2532 GB$0.932/hreastus
7 more instances can run Step 3 VL 10B
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Step 3.7 Flash262K$0.200$1.15Available
Step 3 VL 10BCurrent
Step Image Edit 2Available
Step 1X Edit 1.2Available
Step 1X EditAvailable

Model IDs

step-3-10b-vl
step-3-vl-10b