Skip to main content
MoonshotAI
Kimi-K2.6
Model
MoonshotAI
Kimi-K2.6

Kimi-K2.6 is an open-source, native multimodal agentic model that combines a Mixture-of-Experts (MoE) transformer architecture (1 trillion total parameters, 32B activated) with the MoonViT vision encoder (400M parameters).

Kimi-K2.6

Description:

Kimi-K2.6 is an open-source, native multimodal agentic model that combines a Mixture-of-Experts (MoE) transformer architecture (1 trillion total parameters, 32B activated) with the MoonViT vision encoder (400M parameters). It features 61 layers, 384 experts (8 selected per token), and a 256,000-token context length. It is designed for long-horizon coding, coding-driven design, proactive autonomous execution, multimodal reasoning, visual analysis, and large-scale agent swarm orchestration across text, image, and video inputs. Kimi-K2.6 was developed by Moonshot AI as a part of the Kimi series. This model is ready for commercial/non-commercial use.

Third-Party Community Consideration

This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party's requirements for this application and use case; see link to Non-NVIDIA Kimi-K2.6 Model Card.

License and Terms of Use:

GOVERNING TERMS: The NIM container is governed by the NVIDIA Software License Agreement and the Product-Specific Terms for NVIDIA AI Products; and the use of the model is governed by the NVIDIA Open Model License Agreement. Additional Information: Modified MIT License.

Deployment Geography:

Global

Use Case:

Developers and enterprises building multimodal AI agents for scenario-specific automation, visual analysis applications, advanced web development with autonomous image search and layout iteration, coding assistance, tool-augmented agentic workflows, and large-scale agent swarm orchestration.

Release Date:

NGC: 04/21/2026 via Kimi-K2.6 on NGC
HuggingFace: 04/20/2026 via Kimi-K2.6 on HuggingFace

Reference(s):

Model Architecture:

Architecture Type: Transformer Network Architecture: Mixture-of-Experts (MoE) with 384 experts (8 routed + 1 shared per token) This model was developed based on moonshotai/Kimi-K2.6. Number of model parameters: 1T (32B activated)

Input:

Input Type(s): Text, Image, Video Input Format(s): String Input Parameters: One-Dimensional (1D) Other Properties Related to Input: The model supports up to 256,000 tokens of context length and accepts multimodal inputs (text, image, video) through the OpenAI-compatible chat endpoint. The MoonViT vision encoder (400M parameters) processes visual inputs natively. The tokenizer uses a vocabulary size of 129,280 tokens.

Output:

Output Type(s): Text Output Format: String Output Parameters: One-Dimensional (1D) Other Properties Related to Output: The model generates text responses based on multimodal inputs including reasoning, analysis, and code generation. Supports both Thinking mode (with reasoning traces) and Instant mode. Configurable sampling parameters (temperature, top_p, top_k) and streaming output via the OpenAI-compatible API.

Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.

Software Integration:

Runtime Engine(s): SGLang Supported Hardware Microarchitecture Compatibility:

  • NVIDIA Blackwell
  • NVIDIA Hopper

Supported Operating System(s): Linux

The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.

Training, Testing, and Evaluation Datasets:

Training Dataset:

Data Modality:

  • Text
  • Image
  • Video Image Training Data Size: Undisclosed Text Training Data Size: [More than 10 Trillion Tokens] Video Training Data Size: Undisclosed Data Collection Method by dataset: [Hybrid: Automated, Human, Synthetic] Labeling Method by dataset: [Hybrid: Human, Synthetic]

Testing Dataset:

Data Collection Method by dataset: [Unknown] Labeling Method by dataset: [Unknown]

Evaluation Dataset:

Data Collection Method by dataset: [Hybrid: Automated, Human] Labeling Method by dataset: [Hybrid: Automated, Human] Properties (Quantity, Dataset Descriptions, Sensor(s)): Evaluation covers agentic benchmarks (HLE, BrowseComp, DeepSearchQA, APEX-Agents), reasoning and knowledge (AIME 2026, HMMT 2026, IMO-AnswerBench, GPQA-Diamond), coding (SWE-Bench Verified, Terminal-Bench 2.0, LiveCodeBench), and vision assessments (MMMU-Pro, MathVision, CharXiv).

BenchmarkKimi K2.6
Agentic
HLE-Full (w/ tools)54.0
BrowseComp83.2
BrowseComp (Agent Swarm)86.3
DeepSearchQA (f1-score)92.5
DeepSearchQA (accuracy)83.0
WideSearch (item-f1)80.8
Toolathlon50.0
MCPMark55.9
APEX-Agents27.9
OSWorld-Verified73.1
Reasoning & Knowledge
HLE-Full34.7
AIME 202696.4
HMMT 2026 (Feb)92.7
IMO-AnswerBench86.0
GPQA-Diamond90.5
Vision
MMMU-Pro79.4
MMMU-Pro (w/ python)80.1
CharXiv (RQ)80.4
MathVision87.4
MathVision (w/ python)93.2
BabyVision39.8
Coding
Terminal-Bench 2.066.7
SWE-Bench Pro58.6
SWE-Bench Multilingual76.7
SWE-Bench Verified80.2
SciCode52.2
OJBench (python)60.6
LiveCodeBench (v6)89.6

Inference:

Acceleration Engine: SGLang Test Hardware:

  • NVIDIA B200
  • NVIDIA H200

Ethical Considerations:

NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.

Please make sure you have proper rights and permissions for all input image and video content; if image or video includes people, personal health information, or intellectual property, the image or video generated will not blur or maintain proportions of image subjects included.

Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.

Get Help

NVIDIA Developer Community Forum

Get access to community knowledge base articles and support cases (https://forums.developer.nvidia.com/)

Publisher
MoonshotAI
LicenseNVIDIA proprietary
Latest Versionh200-config-v5
UpdatedAugust 10, 2026 UTC
Compressed Size873 B