Qwen3.8 Flash is Alibaba's language model with a 1.0M context window and up to 131K output tokens, available from 6 providers, starting at $0.09 / 1M input and $0.282 / 1M output. A multimodal mixture-of-experts LLM previewing Alibaba's next-generation architecture, activating only a fraction of its large parameter count per token for efficient reasoning and tool use.
Capabilities
Input4/5
Text✓
Image✓
Audio·
Video✓
PDF✓
Output1/5
Text✓
Image·
Audio·
Video·
Embedding·
Capabilities6/13
Reasoning✓
Adaptive Reasoning·
Function Calling✓
Parallel Function Calling·
Structured Outputs✓
Native JSON Schema✓
Web Search✓
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching✓
Assistant Prefill·
Pricing by Provider
US Dollar ($)
Per 1M tokens
| Provider | Standard | Batch | ||||
|---|---|---|---|---|---|---|
| Input $ / 1M | Output $ / 1M | Cache Read $ / 1M | Cache Write 5m $ / 1M | Input $ / 1M | Output $ / 1M | |
| $0.113 | $0.38 | $0.0141 | $0.176 | — | — | |
| $0.15 | $0.47 | $0.016 | $0.2 | $0.075 | $0.235 | |
| $0.15 | $0.47 | $0.016 | $0.2 | — | — | |
| $0.15 | $0.47 | $0.016 | $0.2 | — | — | |
| $0.09 | $0.282 | N/A | N/A | — | — | |
| $0.15 | $0.47 | $0.016 | $0.2 | — | — | |
Cost Calculator
US Dollar ($)
Preset:
Cheapest Instances to Run It
Cloud GPU instances that can host Qwen3.8 Flash, ranked by cheapest on-demand price. The model needs about 432 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.
All clouds
FP16 (full precision)
US Dollar ($)
Instance | Cloud | GPU | VRAM | Price | Cheapest region | |
|---|---|---|---|---|---|---|
| a3-megagpu-8g | 8× nvidia-h100-mega-80gb | 640 GB | $10.42/hr | |||
| a3-ultragpu-8g | 8× nvidia-h200-141gb | 1128 GB | $13.82/hr | |||
| p4de.24xlarge | 8× A100 | 640 GB | $27.45/hr | |||
Versions
| Version | Released | Context | Input / 1M | Output / 1M | Status |
|---|---|---|---|---|---|
| Qwen Audio 3 Flash ASR | — | — | — | — | Available |
| Qwen Audio 3 Flash ASR Filetrans | — | — | — | — | Available |
| Qwen Audio 3 Flash ASR Streaming | — | — | — | — | Available |
| Qwen Audio 3 Flash Realtime | — | — | $0.230 | $0.930 | Available |
| Qwen3.8 Flash | 1.0M | $0.090 | $0.282 | Current | |
| Qwen3.8-Flash-Next | 1.0M | — | — | Available | |
| Qwen3.5 Omni Flash | — | — | $0.400 | $3.00 | Available |
| Qwen Flash Character | — | — | $0.050 | $0.400 | Available |
| Qwen TTS Flash | — | — | $0.230 | $1.43 | Available |
Other Models
| Model | Tier | Released | Context | Input / 1M | Output / 1M |
|---|---|---|---|---|---|
| Qwen3.8 2.4T A95B | — | 1.0M | $2.00 | $6.00 | |
| Qwen Image 3 | — | — | — | — | — |
| Qwen Audio 3 Plus Realtime | Plus | — | — | $0.800 | $6.40 |
| Qwen Audio 3 Plus TTS | Plus | — | — | — | — |
| Qwen Audio 3.1 Plus Realtime | Plus | — | — | $0.800 | $6.40 |
| Qwen Image 3.0 Pro | Pro | — | — | — | — |
| EAGLE Qwen 2.5 3B Instruct | — | — | — | — | — |
| DeepSeek R1 Qwen3 8B | — | — | — | — | — |
| Qwen3.8 Max | Max | 1.0M | $1.65 | $4.95 | |
| Qwen3.8 27B | — | 1.0M | $0.400 | $1.49 |