deepseek-ai
DeepSeek-V4-Flash
Model
deepseek-ai
DeepSeek-V4-Flash

DeepSeek-V4-Flash is a Mixture-of-Experts (MoE) language model with 284B total parameters and 13B activated parameters. DeepSeek-V4-Flash was developed by DeepSeek as a part of DeepSeek-V4 collection.

Sign in to access all content for this ModelSigning in will also allow download accessSign In

DeepSeek-V4-Flash

Description:

DeepSeek-V4-Flash is a Mixture-of-Experts (MoE) language model with 284B total parameters and 13B activated parameters. DeepSeek-V4-Flash was developed by DeepSeek as a part of DeepSeek-V4 collection. This model is ready for commercial/non-commercial use.

Third-Party Community Consideration:

This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party's requirements for this application and use case; see link to Non-NVIDIA DeepSeek-V4-Flash Model Card.

License/Terms of Use:

GOVERNING TERMS: The NIM container is governed by the NVIDIA Software and Model Evaluation license; Use of this model is governed by the NVIDIA Open Model Agreement.

Additional Information: MIT License.

You are responsible for ensuring that your use of NVIDIA provided models complies with all applicable laws.

Deployment Geography:

Global

Use Case:

DeepSeek V4 is well-suited for advanced reasoning, agentic AI applications, tool use scenarios, and complex problem-solving in domains such as mathematics, software engineering, and enterprise AI assistants.

Release Date:

NGC 04/27/2026 via DeepSeek-V4-Flash on NGC

Reference(s):

References:

Model Architecture:

Architecture Type: Transformer Network Architecture: Mixture of Experts (MoE) with Hybrid Attention (Compressed Sparse Attention + Heavily Compressed Attention) Total Parameters: 284B Active Parameters: 13B

Input:

Input Types: Text Input Formats: String Input Parameters: One Dimensional (1D) Other Input Properties: Supports multi-turn conversations with system prompts, user messages, and assistant responses. Maximum context length of 1 million tokens. Uses a custom encoding pipeline (encoding_dsv4) with three reasoning modes: Non-think (fast), Think High (logical analysis), and Think Max (full reasoning extent).

Output:

Output Types: Text Output Format: String Output Parameters: One Dimensional (1D) Other Output Properties: Supports structured JSON output, function/tool calling, and reasoning content when enabled.

Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.

Software Integration:

Runtime Engines:

  • SGLang

Supported Hardware:

  • NVIDIA Blackwell: NVIDIA B200 Tensor Core GPU
  • NVIDIA Hopper: NVIDIA H100 Tensor Core GPU, NVIDIA H200 Tensor Core GPU, NVIDIA H20 Tensor Core GPU

Operating Systems: Linux

The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.

Model Version(s):

DeepSeek-V4-Flash

Training, Testing, and Evaluation Datasets:

Training Dataset

Data Modality: Text Text Training Data Size: [More than 10 Trillion Tokens] Data Collection Method by dataset: Undisclosed Labeling Method by dataset: Undisclosed Properties (Quantity, Dataset Descriptions, Sensor(s)): Two-stage post-training pipeline: (1) independent cultivation of domain-specific experts via SFT and RL with GRPO, (2) unified model consolidation via on-policy distillation. Uses Muon optimizer for faster convergence and training stability.

Testing Dataset

Data Collection Method by dataset: Undisclosed Labeling Method by dataset: Undisclosed Properties (Quantity, Dataset Descriptions, Sensor(s)): Undisclosed

Evaluation Dataset

Benchmark Score:

Benchmark (Metric)V4-Flash Non-ThinkV4-Flash HighV4-Flash MaxV4-Pro Non-ThinkV4-Pro HighV4-Pro Max
Knowledge & Reasoning
MMLU-Pro (EM)83.086.486.282.987.187.5
SimpleQA-Verified (Pass@1)23.128.934.145.046.257.9
Chinese-SimpleQA (Pass@1)71.573.278.975.877.784.4
GPQA Diamond (Pass@1)71.287.488.172.989.190.1
HLE (Pass@1)8.129.434.87.734.537.7
LiveCodeBench (Pass@1)55.288.491.656.889.893.5
Codeforces (Rating)-28163052-29193206
HMMT 2026 Feb (Pass@1)40.891.994.831.794.095.2
IMOAnswerBench (Pass@1)41.985.188.435.388.089.8
Apex (Pass@1)1.019.133.00.427.438.3
Apex Shortlist (Pass@1)9.372.185.79.285.590.2
Long Context
MRCR 1M (MMR)37.576.978.744.783.383.5
CorpusQA 1M (ACC)15.559.360.535.656.562.0
Agentic
Terminal Bench 2.0 (Acc)49.156.656.959.163.367.9
SWE Verified (Resolved)73.778.679.073.679.480.6
SWE Pro (Resolved)49.152.352.652.154.455.4
SWE Multilingual (Resolved)69.770.273.369.874.176.2
BrowseComp (Pass@1)-53.573.2-80.483.4
HLE w/ tools (Pass@1)-40.345.1-44.748.2
MCPAtlas (Pass@1)64.067.469.069.474.273.6
GDPval-AA (Elo)--1395--1554
Toolathlon (Pass@1)40.743.547.846.349.051.8

Data Collection Method by dataset: [Automated] Labeling Method by dataset: [Human] Properties (Quantity, Dataset Descriptions, Sensor(s)): Evaluated on competitive programming, mathematical reasoning, and general reasoning benchmarks.

Inference

Acceleration Engine: SGLang Test Hardware:

  • NVIDIA B200 Tensor Core GPU
  • NVIDIA H100 Tensor Core GPU
  • NVIDIA H200 Tensor Core GPU
  • NVIDIA H20 Tensor Core GPU

Precision formats: FP8

Ethical Considerations

NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.

Users are responsible for model inputs and outputs. Users are responsible for ensuring safe integration of this model, including implementing guardrails as well as other safety mechanisms, prior to deployment.

Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.

Get Help

Getting started with the NIM

Deploying and integrating the NIM is straightforward thanks to our industry standard APIs. Visit the NIM Container page for release documentation, deployment guides and more.

NVIDIA Developer Community Forum

Get access to community knowledge base articles and support cases (NVIDIA Developer Forums).

Publisher
deepseek-ai
Latest Versionhf-nim-fp8-checksum
UpdatedMay 13, 2026 UTC
Compressed Size273.86 GB