Qwen3 VL 8B Instruct is Alibaba's language model with a 262K context window and up to 33K output tokens, available from 5 providers, starting at $0.08 / 1M input and $0.2 / 1M output. An instruction-tuned 8B vision-language model from the Qwen3 series, optimized for conversational multimodal tasks involving text and image inputs.
Specifications
Canonical IDalibaba-qwen3-vl-8b-instruct
TypeLanguage
StatusDeprecated
CreatorAlibabaAlibaba
Providers
Context Window262K tokens
Max Output33K tokens
Input ModalitiesImageText
Output ModalitiesText
Parameters8B
HuggingFace Likes874
HuggingFace Downloads (30d)3,765,920
HuggingFace Downloads (all-time)23,111,974
Release Date · 11 months ago
Deprecation Date
Benchmarks
Intelligence Index
#403
Math Index
#204
MMLU-Pro
#235
GPQA
#434
HLE
#529
LiveCodeBench
#210
IFBench
#341
Time to First Token
#476
SciCode
#441
AIME 2025
#204
LCR
#358
TerminalBench Hard
#321
TAU2
#264
Output TPS
#117

Capabilities

Input2/5
Text
Image
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities4/13
Reasoning·
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardBatch
Input
$ / 1M
Output
$ / 1M
Input
$ / 1M
Output
$ / 1M
Alibaba Qwen logo
Alibaba Qwen
qwen3-vl-8b-instruct
$0.18$0.7$0.09$0.35
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/qwen3-vl-8b-instruct
$0.2$0.2
Novita logo
Novita
novita/qwen/qwen3-vl-8b-instruct
$0.08$0.5
OpenRouter logo
OpenRouter
qwen/qwen3-vl-8b-instruct
$0.117$0.455
Together AI logo
Together AI
together_ai/Qwen/Qwen3-VL-8B-Instruct
$0.18$0.68

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Qwen3 VL 8B Instruct, ranked by cheapest on-demand price. The model needs about 19 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g6.xlargeAWSL422 GB$0.805/hr
Standard_NV16as_v4AzureAMD Radeon Instinct MI2532 GB$0.932/hr
g6.2xlargeAWSL422 GB$0.978/hr
7 more instances can run Qwen3 VL 8B Instruct
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Qwen3 VL 32B Instruct262K$0.104$0.416Deprecated
Qwen3 VL 8B Instruct262K$0.080$0.200Current

Model IDs

accounts/fireworks/models/qwen3-vl-8b-instruct
alibaba-qwen3-vl-8b-instruct
fireworks_ai/accounts/fireworks/models/qwen3-vl-8b-instruct
huggingface-vlm-qwen3-vl-8b-instruct
novita/qwen/qwen3-vl-8b-instruct
openrouter/qwen/qwen3-vl-8b-instruct
qwen/qwen3-vl-8b-instruct
qwen3-vl-8b-instruct
together_ai/Qwen/Qwen3-VL-8B-Instruct