Qwen3 VL 30B A3B Instruct is Alibaba's language model with a 262K context window and up to 33K output tokens, available from 5 providers, starting at $0.13 / 1M input and $0.52 / 1M output. An instruction-tuned vision-language MoE model with 30B total and 3B activated parameters, offering strong multimodal understanding and generation capabilities.
Specifications
Canonical IDalibaba-qwen3-vl-30b-a3b-instruct
TypeLanguage
StatusActive
CreatorAlibabaAlibaba
Providers
Context Window262K tokens
Max Output33K tokens
Input ModalitiesImageText
Output ModalitiesText
Parameters30B
HuggingFace Likes562
HuggingFace Downloads (30d)2,219,395
HuggingFace Downloads (all-time)14,070,852
Release Date · 10 months ago
Knowledge Cutoff · 1 year ago
Benchmarks
Intelligence Index
10.0
#296
Math Index
72.3
#85
MMLU-Pro
0.8
#146
GPQA
0.7
#213
HLE
0.1
#231
LiveCodeBench
0.5
#143
IFBench
0.3
#308
Time to First Token
0.94s
#354
SciCode
0.3
#237
AIME 2025
0.7
#85
LCR
0.2
#253
TerminalBench Hard
0.1
#246
TAU2
0.2
#322
Output TPS

Capabilities

Input2/5
Text
Image
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities4/13
Reasoning·
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardBatch
Input
$ / 1M
Output
$ / 1M
Input
$ / 1M
Output
$ / 1M
Alibaba Qwen logo
Alibaba Qwen
qwen3-vl-30b-a3b-instruct
$0.2$0.8$0.1$0.4
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/qwen3-vl-30b-a3b-instruct
$0.15$0.6
Hugging Face logo
Hugging Face
novita:qwen/qwen3-vl-30b-a3b-instruct
$0.2$0.7
Novita logo
Novita
novita/qwen/qwen3-vl-30b-a3b-instruct
$0.2$0.7
OpenRouter logo
OpenRouter
qwen/qwen3-vl-30b-a3b-instruct
$0.13$0.52

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Qwen3 VL 30B A3B Instruct, ranked by cheapest on-demand price. The model needs about 72 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_NP20sAzure2× AMD Alveo U250 FPGA (64GB)128 GB$3.30/hrwestus2
g7e.2xlargeAWSRTX PRO Server 600096 GB$3.36/hrus-east-1
Standard_NC24ads_A100_v4AzureNVIDIA A10080 GB$3.67/hrwestus2
7 more instances can run Qwen3 VL 30B A3B Instruct
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Qwen3 VL 30B A3B Instruct262K$0.130$0.520Current
Qwen3 VL 30B A3B Thinking262K$0.130$0.600Available
Qwen3 VL 235B A22B Instruct262K$0.210$0.880Available
Qwen3 VL 235B A22B Thinking262K$0.220$0.880Available

Model IDs

accounts/fireworks/models/qwen3-vl-30b-a3b-instruct
alibaba-qwen3-vl-30b-a3b-instruct
fireworks_ai/accounts/fireworks/models/qwen3-vl-30b-a3b-instruct
novita/qwen/qwen3-vl-30b-a3b-instruct
qwen/qwen3-vl-30b-a3b-instruct
Qwen/Qwen3-VL-30B-A3B-Instruct
qwen3-vl-30b-a3b-instruct