Skip to main content
NVIDIA
DeepSeek-V4.1-Flash
Model
NVIDIA
DeepSeek-V4.1-Flash

DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens.

DeepSeek-V4.1-Flash

Description

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model for long-context reasoning, coding, agentic workflows, and visual understanding. It accepts text and image inputs and generates text.

This model is ready for commercial or non-commercial use.

Third-Party Community Consideration

This model is not owned or developed by NVIDIA. DeepSeek AI developed and built the model for this application and use case; see link to Non-NVIDIA DeepSeek AI DeepSeek-V4.1-Flash Model Card.

License and Terms of Use

GOVERNING TERMS: Use of this container is governed by the NVIDIA Software License Agreement and Product-Specific Terms for NVIDIA AI Products; and the use of this model is governed by the NVIDIA Open Model Agreement. ADDITIONAL INFORMATION: MIT License.

Deployment Geography

Global

Use Case

Developers and enterprises can use DeepSeek-V4.1-Flash to build multimodal assistants, long-context analysis systems, coding applications, visual understanding workflows, and agentic applications.

Release Date

NGC 09/30/2026 via DeepSeek-V4.1-Flash model on NGC

Reference(s)

Model Architecture

Architecture Type: Transformer

Network Architecture: A 40-layer causal encoder-decoder multimodal Mixture-of-Experts architecture with a 20-layer causal encoder and a 20-layer decoder.

Total Parameters: 552B backbone parameters, plus 196B sparsely accessed Engram conditional-memory parameters

Active Parameters: 8B per token during prefill and 16B per token during decode

Input

Input Types: Text, Image

Input Formats: Text: String; Image: Red, Green, Blue (RGB)

Text Input Parameters: One-Dimensional (1D)

Image Input Parameters: Two-Dimensional (2D)

Other Input Properties: Text prompts, chat messages, image inputs, and interleaved image-and-text inputs. The model supports context lengths up to 1M tokens.

Output

Output Types: Text

Output Format: String

Output Parameters: One-Dimensional (1D)

Other Output Properties: Generated natural language, structured text, code, and image-grounded answers.

Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.

Software Integration

Runtime Engines:

  • SGLang

Operating Systems: Linux

Supported Hardware: NVIDIA B200 Tensor Core GPU, NVIDIA H20 Tensor Core GPU, NVIDIA H200 Tensor Core GPU

The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.

Model Version(s)

DeepSeek-V4.1-Flash 2.1.4-variant. The model can be integrated into AI systems through the NIM's OpenAI-compatible APIs.

Training, Testing, and Evaluation Datasets

Training Dataset

Data Modality: Text and image

Text Training Data Size: [More than 10 Trillion Tokens]

Image Training Data Size: Undisclosed

Data Collection Method by dataset: Undisclosed

Labeling Method by dataset: Undisclosed

Properties: Data count: 45T tokens; modalities: text and image; content nature: Undisclosed; linguistic characteristics: Undisclosed.

Testing Dataset

Data Collection Method by dataset: Undisclosed

Labeling Method by dataset: Undisclosed

Properties: Data count: Undisclosed; modalities: text and image; content nature: Undisclosed; linguistic characteristics: Undisclosed.

Evaluation Dataset

Data Collection Method by dataset: Undisclosed

Labeling Method by dataset: Undisclosed

Properties: Data count: Undisclosed; modalities: text and image; content nature: Undisclosed; linguistic characteristics: Undisclosed.

Inference

Runtime: SGLang

Acceleration Engine: SGLang

Precision: 8-bit floating-point (FP8)

Test Hardware:

  • NVIDIA B200 Tensor Core GPU
  • NVIDIA H20 Tensor Core GPU
  • NVIDIA H200 Tensor Core GPU

Ethical Considerations

NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal developer team to ensure this model meets requirements for the relevant industry and use case and address unforeseen product misuse.

Users are responsible for model inputs and outputs. Users are responsible for ensuring safe integration of this model, including implementing guardrails as well as other safety mechanisms, prior to deployment.

Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.

Please make sure you have proper rights and permissions for all input image and video content; if image or video includes people, personal health information, or intellectual property, the image or video generated will not blur or maintain proportions of image subjects included.

Get Help

NVIDIA Developer Community Forum

Get access to community knowledge base articles and support cases through the NVIDIA Developer Forums.

Publisher
NVIDIA
LicenseNVIDIA proprietary
Latest Versionhf-dba1be0-nim-v2
UpdatedSeptember 30, 2026 UTC
Compressed Size475.26 GB

Modal Content

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.