Qwen3.5 9B is Alibaba's language model with a 262K context window and up to 236K output tokens, available from 4 providers, starting at $0.1 / 1M input and $0.15 / 1M output. A 9B-parameter dense LLM from the Qwen3.5 series, offering a strong balance of capability and efficiency for general text generation tasks.
Specifications
Canonical IDalibaba-qwen3-5-9b
TypeLanguage
StatusActive
CreatorAlibabaAlibaba
Providers
Context Window262K tokens
Max Output236K tokens
Input ModalitiesImageTextVideo
Output ModalitiesText
Reasoning Effortsdefault
Parameters9B
HuggingFace Likes1,315
HuggingFace Downloads (30d)6,481,835
HuggingFace Downloads (all-time)9,617,908
Release Date · 6 months ago
Benchmarks
Intelligence Index
#192
Coding Index
#116
GPQA
#143
HLE
#147
IFBench
#87
Time to First Token
#343
SciCode
#332
LCR
#136
TerminalBench Hard
#133
TAU2
#84
Output TPS
#129

Capabilities

Input3/5
Text
Image
Audio·
Video
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities5/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardBatch
Input
$ / 1M
Output
$ / 1M
Input
$ / 1M
Output
$ / 1M
DeepInfra logo
DeepInfra
deepinfra/Qwen/Qwen3.5-9B
$0.1$0.15
Hugging Face logo
Hugging Face
ovhcloud:Qwen3.5-9B
$0.12$0.18
Hugging Face logo
Hugging Face
together_ai:Qwen/Qwen3.5-9B
$0.17$0.25
OpenRouter logo
OpenRouter
qwen/qwen3.5-9b
$0.1$0.15N/AN/A
OpenRouter logo
OpenRouter
qwen/qwen3.5-9b:batch
$0.17$0.25
Together AI logo
Together AI
together_ai/Qwen/Qwen3.5-9B
$0.17$0.25

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Qwen3.5 9B, ranked by cheapest on-demand price. The model needs about 22 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g2-standard-4GCPnvidia-l424 GB$0.705/hr
g6.xlargeAWSL422 GB$0.805/hr
g2-standard-8GCPnvidia-l424 GB$0.851/hr
7 more instances can run Qwen3.5 9B
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Qwen3.5 9B262K$0.100$0.150Current
Qwen3.5 27B262K$0.195$1.56Available

Model IDs

accounts/fireworks/models/qwen3p5-9b
alibaba-qwen3-5-9b
deepinfra/Qwen/Qwen3.5-9B
huggingface-vlm-qwen3-5-9b
qwen/qwen3.5-9b
Qwen/Qwen3.5-9B
qwen3-5-9b
qwen3-5-9b-non-reasoning
together_ai/Qwen/Qwen3.5-9B