JT-4.1 Flash 236B A21B Non-Reasoning is Jt's language model. A large mixture-of-experts LLM from China Mobile with 236B total and 21B active parameters, tuned for fast non-reasoning responses.
| Benchmarks | |
|---|---|
| Intelligence Index | #97 |
| Coding Index | #64 |
| GPQA | #110 |
| HLE | #154 |
| Time to First Token | #172 |
| LCR | #131 |
| Output TPS | #399 |
Capabilities
Input1/5
Text✓
Image·
Audio·
Video·
PDF·
Output1/5
Text✓
Image·
Audio·
Video·
Embedding·
Capabilities0/13
Reasoning·
Adaptive Reasoning·
Function Calling·
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·
Cheapest Instances to Run It
Cloud GPU instances that can host JT-4.1 Flash 236B A21B Non-Reasoning, ranked by cheapest on-demand price. The model needs about 566 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.
All clouds
FP16 (full precision)
US Dollar ($)
Instance | Cloud | GPU | VRAM | Price | Cheapest region | |
|---|---|---|---|---|---|---|
| p4de.24xlarge | 8× A100 | 640 GB | $27.45/hr | |||
| Standard_ND96amsr_A100_v4 | 8× NVIDIA A100 (80GB) | 640 GB | $32.77/hr | |||
| g7e.48xlarge | 8× RTX PRO Server 6000 | 768 GB | $33.14/hr | |||
Versions
| Version | Released | Context | Input / 1M | Output / 1M | Status |
|---|---|---|---|---|---|
| JT-4.1 Flash 236B A21B Non-Reasoning | — | — | — | — | Current |
| JT Flash 35B | — | — | — | — | Available |
Other Models
| Model | Tier | Released | Context | Input / 1M | Output / 1M |
|---|---|---|---|---|---|
| JT-35B-Flash | Mini | — | — | — | — |
| JT-MINI | Mini | — | — | — | — |