- Chat
Kimi K3
Open-weight mixture-of-experts model: 2.8T total, 104B active parameters.
Kimi model family
About this model
Kimi K3 is a 2.8T total, 104B active multimodal mixture-of-experts model. Its model card describes 93 layers, Kimi Delta Attention with Gated MLA, 896 experts with 16 selected per token and 2 shared experts, a MoonViT-V2 vision encoder, a 1,048,576-token context window, and MXFP4 weights with MXFP8 activations from quantization-aware training. The lineup is subject to change; a successor may ship in its place.
- Parameters
- 2.8T104B active per token
- Native context
- 1,048,5761,048,576 tokens served by Arnict.
- Output guidance
- —Published guidance
Capabilities
- Chat completionsUse this model with the chat-completions API when it is marked Available.
- Native input modalitiesText input is listed in the published catalog metadata.
- Long-context input1,048,576 tokens · 1,048,576 tokens served by Arnict.
Published limits
Published model-card facts, kept separate from live availability.
| Metric | Published value | Note |
|---|---|---|
| Arnict API output limit | Not published | Maximum output accepted by the Arnict API for this model |
| Parameters | 2.8T total / 104B active | Upstream model card |
| Context window | 1,048,576 tokens | 1,048,576 tokens served by Arnict. |
| Output guidance | Not published by upstream | Upstream model card |
| Weights / activations | MXFP4 weights / MXFP8 activations (quantization-aware training) | Published format |
- 93 layers
- Kimi Delta Attention (KDA) + Gated MLA
- 896 experts; 16 selected per token; 2 shared experts
- MoonViT-V2 vision encoder (401M parameters)
- Native text and image modality
Use moonshotai/kimi-k3 with an OpenAI-compatible client after this model is enabled.
Related models
GLM
GLM 5.3 Flash Abliterated
A derivative of GLM 5.3 Flash with modified refusal behavior for general-purpose inference.
GLM
GLM 5.3 Abliterated
A derivative of GLM 5.3 with modified refusal behavior for general-purpose inference.
DeepSeek
DeepSeek V4.1 Flash
Multimodal mixture-of-experts model: 552B backbone parameters, with 8B active during prefill and 16B during decoding.
