Models / Qwen
  • Chat
  • Reasoning

Qwen 3.8 2.4T A95B

Text-only mixture-of-experts model: 2.4T total, 95B active parameters.

Qwen model family

About this model

Qwen 3.8 2.4T A95B is a text-only 2.4T total, 95B active mixture-of-experts causal language model. Its model card lists 92 layers, 512 experts with 10 routed and 1 shared expert activated, Gated DeltaNet with Gated Attention, a 262,144-token native context, and output guidance of 262,144 reasoning tokens or 131,072 final-response tokens; thinking is required. The lineup is subject to change; a successor may ship in its place.

Parameters
2.4T95B active per token
Native context
262,144262,144 tokens served by Arnict.
Output guidance
262,144Published guidance

Capabilities

  • Chat completionsUse this model with the chat-completions API when it is marked Available.
  • Native input modalitiesText input is listed in the published catalog metadata.
  • Long-context input262,144 tokens · 262,144 tokens served by Arnict.
  • Published output guidance262,144 tokens · Reasoning: 262,144 · final response: 131,072

Published limits

Published model-card facts, kept separate from live availability.

MetricPublished valueNote
Arnict API output limit131,072 tokensMaximum output accepted by the Arnict API for this model
Parameters2.4T total / 95B activeUpstream model card
Context window262,144 tokens262,144 tokens served by Arnict.
Output guidance262,144 tokens recommended maximumReasoning: 262,144 · final response: 131,072
Weights / activationsBF16Published format

  • 92 layers
  • 512 experts; 10 routed + 1 shared activated
  • Gated DeltaNet + Gated Attention
  • Text-only; thinking required

Use qwen/qwen3.8-2.4t-a95b with an OpenAI-compatible client after this model is enabled.

Read the API docs