- Chat
- Reasoning
Qwen 3.8 2.4T A95B
Text-only mixture-of-experts model: 2.4T total, 95B active parameters.
Qwen model family
About this model
Qwen 3.8 2.4T A95B is a text-only 2.4T total, 95B active mixture-of-experts causal language model. Its model card lists 92 layers, 512 experts with 10 routed and 1 shared expert activated, Gated DeltaNet with Gated Attention, a 262,144-token native context, and output guidance of 262,144 reasoning tokens or 131,072 final-response tokens; thinking is required. The lineup is subject to change; a successor may ship in its place.
- Parameters
- 2.4T95B active per token
- Native context
- 262,144262,144 tokens served by Arnict.
- Output guidance
- 262,144Published guidance
Capabilities
- Chat completionsUse this model with the chat-completions API when it is marked Available.
- Native input modalitiesText input is listed in the published catalog metadata.
- Long-context input262,144 tokens · 262,144 tokens served by Arnict.
- Published output guidance262,144 tokens · Reasoning: 262,144 · final response: 131,072
Published limits
Published model-card facts, kept separate from live availability.
| Metric | Published value | Note |
|---|---|---|
| Arnict API output limit | 131,072 tokens | Maximum output accepted by the Arnict API for this model |
| Parameters | 2.4T total / 95B active | Upstream model card |
| Context window | 262,144 tokens | 262,144 tokens served by Arnict. |
| Output guidance | 262,144 tokens recommended maximum | Reasoning: 262,144 · final response: 131,072 |
| Weights / activations | BF16 | Published format |
- 92 layers
- 512 experts; 10 routed + 1 shared activated
- Gated DeltaNet + Gated Attention
- Text-only; thinking required
Use qwen/qwen3.8-2.4t-a95b with an OpenAI-compatible client after this model is enabled.
Related models
Qwen
Qwen 3.6 35B A3B
Vision-language mixture-of-experts model: 35B total, 3B active parameters.
GLM
GLM 5.3 Flash
Long-context chat, coding and tool workflows with GLM 5.3 Flash.
DeepSeek
DeepSeek V4.1 Flash
Multimodal mixture-of-experts model: 552B backbone parameters, with 8B active during prefill and 16B during decoding.
