Qwen3.8 2.4T A95B is Alibaba's language model with a 1.0M context window and up to 262K output tokens, available from 6 providers, starting at $2.00 / 1M input and $6.00 / 1M output. A sparse mixture-of-experts LLM with 95 billion active parameters out of 2.4 trillion total, serving as the open-weight counterpart to Qwen3.8 Max.
Specifications
Canonical IDalibaba-qwen3-8-2-4t-a95b
TypeLanguage
StatusActive
CreatorAlibabaAlibaba
Providers
Context Window1.0M tokens
Max Output262K tokens
Input ModalitiesImageText
Output ModalitiesText
Reasoning Effortsdefault
Parameters2.45T
HuggingFace Likes625
HuggingFace Downloads (30d)978
HuggingFace Downloads (all-time)978
Release Date · 22 days ago
Benchmarks
Intelligence Index
#9
Coding Index
#16
GPQA
#7
HLE
#19
Time to First Token
#492
SciCode
#31
LCR
#39
Output TPS
#217

Capabilities

Input2/5
Text
Image
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities5/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling·
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardBatch
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Alibaba Qwen logo
Alibaba Qwen
qwen3.8-2.4t-a95b
$2.00$6.00N/A$1.00$3.00N/A
DeepInfra logo
DeepInfra
deepinfra/Qwen/Qwen3.8-2.4T-A95B
$2.00$6.00$0.2
Hugging Face logo
Hugging Face
novita:qwen/qwen3.8-2.4t-a95b
$2.00$6.00N/A
OpenRouter logo
OpenRouter
qwen/qwen3.8-2.4t-a95b
$2.00$6.00$0.2N/AN/AN/A
OpenRouter logo
OpenRouter
qwen/qwen3.8-2.4t-a95b:batch
$2.00$6.00$0.25
Together AI logo
Together AI
together_ai/Qwen/Qwen3.8-2.4T-A95B
$2.50$6.25$0.5
Vercel AI Gateway logo
Vercel AI Gateway
alibaba/qwen3.8-2.4t-a95b
$2.00$6.00$0.25

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Qwen3.8 2.4T A95B, ranked by cheapest on-demand price. The model needs about 5871 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)

No GPU instances fit Qwen3.8 2.4T A95B at FP16 precision — try a lower precision such as INT4.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Qwen3.8 2.4T A95B1.0M$2.00$6.00Current
Qwen Audio 3 Flash ASRAvailable
Qwen Audio 3 Flash ASR FiletransAvailable
Qwen Audio 3 Flash ASR StreamingAvailable
Qwen Audio 3 Flash Realtime$0.450$4.50Available
Qwen Audio 3 Plus Realtime$0.800$6.40Available
Qwen Audio 3 Plus TTSAvailable
Qwen Image 3Available
Qwen Image 3.0 ProAvailable
EAGLE Qwen 2.5 3B InstructAvailable
DeepSeek R1 Qwen3 8BAvailable

Model IDs

accounts/fireworks/models/qwen3p8-2p4t-a95b
alibaba-qwen3-8-2-4t-a95b
alibaba/qwen3.8-2.4t-a95b
deepinfra/Qwen/Qwen3.8-2.4T-A95B
qwen/qwen3.8-2.4t-a95b
Qwen/Qwen3.8-2.4T-A95B
qwen3-8-2-4t-a95b
qwen3.8-2.4t-a95b
together_ai/Qwen/Qwen3.8-2.4T-A95B