Models / Kimi
  • Chat

Kimi K3

Open-weight mixture-of-experts model: 2.8T total, 104B active parameters.

Kimi model family

About this model

Kimi K3 is a 2.8T total, 104B active multimodal mixture-of-experts model. Its model card describes 93 layers, Kimi Delta Attention with Gated MLA, 896 experts with 16 selected per token and 2 shared experts, a MoonViT-V2 vision encoder, a 1,048,576-token context window, and MXFP4 weights with MXFP8 activations from quantization-aware training. The lineup is subject to change; a successor may ship in its place.

Parameters
2.8T104B active per token
Native context
1,048,5761,048,576 tokens served by Arnict.
Output guidance
—Published guidance

Capabilities

  • Chat completionsUse this model with the chat-completions API when it is marked Available.
  • Native input modalitiesText input is listed in the published catalog metadata.
  • Long-context input1,048,576 tokens · 1,048,576 tokens served by Arnict.

Published limits

Published model-card facts, kept separate from live availability.

MetricPublished valueNote
Arnict API output limitNot publishedMaximum output accepted by the Arnict API for this model
Parameters2.8T total / 104B activeUpstream model card
Context window1,048,576 tokens1,048,576 tokens served by Arnict.
Output guidanceNot published by upstreamUpstream model card
Weights / activationsMXFP4 weights / MXFP8 activations (quantization-aware training)Published format

  • 93 layers
  • Kimi Delta Attention (KDA) + Gated MLA
  • 896 experts; 16 selected per token; 2 shared experts
  • MoonViT-V2 vision encoder (401M parameters)
  • Native text and image modality

Use moonshotai/kimi-k3 with an OpenAI-compatible client after this model is enabled.

Read the API docs →