Skip to main content
z-AI
GLM-5.3 Flash
Model
z-AI
GLM-5.3 Flash

GLM-5.3-Flash is a natively multimodal mixture-of-experts model.

GLM-5.3-Flash

Description

GLM-5.3-Flash is a natively multimodal mixture-of-experts model with 320B total parameters and 18B active parameters. It uses a hybrid architecture combining sparse and linear attention and Manifold-Constrained Hyper-Connections for coding, reasoning, agentic, and multimodal workloads.

This model is ready for commercial or non-commercial use.

Third-Party Community Consideration

This model is not owned or developed by NVIDIA. It was developed by Z.ai; see link to Non-NVIDIA GLM-5.3-Flash Model Card.

License and Terms of Use

GOVERNING DOWNLOAD TERMS: Use of the model is governed by the NVIDIA Open Model Agreement. ADDITIONAL INFORMATION: MIT License.

Deployment Geography

Global

Use Case

Developers can use GLM-5.3-Flash for coding, reasoning, agentic workflows, long-context text processing, and image or video understanding.

Release Date

NGC 08/31/2026 via GLM-5.3-Flash on NGC

Reference(s)

Model Architecture

Architecture Type: Transformer Network Architecture: Mixture-of-Experts with hybrid sparse and linear attention and Manifold-Constrained Hyper-Connections Total Parameters: 320B Active Parameters: 18B Vocabulary Size: 154,880

Input

Input Types: Text, Image, Video Input Formats: Text: String; Image: Red, Green, Blue (RGB); Video: mp4, mov, webm Input Parameters: Text: One-Dimensional (1D); Image: Two-Dimensional (2D); Video: Three-Dimensional (3D) Other Input Properties: Supports configurable reasoning effort. Input Context Length (ISL): Up to 1,048,576 tokens

Output

Output Types: Text Output Format: String Output Parameters: One-Dimensional (1D) Other Output Properties: Autoregressive text generation.

Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.

Software Integration

Runtime Engines:

  • vLLM

Supported Hardware:

  • NVIDIA Blackwell - B200 Tensor Core GPU
  • NVIDIA Hopper - H20 Tensor Core GPU
  • NVIDIA Hopper - H200 Tensor Core GPU

Precision: FP8

Supported Operating System(s): Linux

The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.

Model Version(s)

GLM-5.3-Flash

Training, Testing, and Evaluation Datasets

Training Dataset

Data Modality: Text, Image, Video Text Training Data Size: [More than 10 Trillion Tokens] Image Training Data Size: Undisclosed Video Training Data Size: Undisclosed Data Collection Method by dataset: Undisclosed Labeling Method by dataset: Undisclosed Properties: A 30T-token multimodal pre-training corpus.

Testing Dataset

Data Collection Method by dataset: Undisclosed Labeling Method by dataset: Undisclosed Properties: Undisclosed

Evaluation Dataset

Data Collection Method by dataset: [Hybrid: Automated, Manually-Collected] Labeling Method by dataset: [Hybrid: Automated, Manually-Labeled] Properties: Evaluated on coding, agentic, reasoning, long-context, and multimodal benchmarks including Terminal-Bench 2.1, DeepSWE, HLE, NL2Repo, and BabyVision.

Inference

Acceleration Engine: vLLM

Test Hardware:

  • NVIDIA B200 Tensor Core GPU
  • NVIDIA H20 Tensor Core GPU
  • NVIDIA H200 Tensor Core GPU

Ethical Considerations

NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal developer team to ensure this model meets requirements for the relevant industry and use case and address unforeseen product misuse.

Please make sure you have proper rights and permissions for all input image and video content; if image or video includes people, personal health information, or intellectual property, the image or video generated will not blur or maintain proportions of image subjects included.

Users are responsible for model inputs and outputs. Users are responsible for ensuring safe integration of this model, including implementing guardrails as well as other safety mechanisms, prior to deployment.

Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.

Get Help

Getting started with the NIM

Deploying and integrating the NIM is straightforward thanks to our industry standard APIs. Visit the NIM documentation for release documentation, deployment guides and more.

NVIDIA Developer Community Forum

Get access to community knowledge base articles and support cases (NVIDIA Developer Forums).

Publisher
z-AI
LicenseNVIDIA proprietary
Latest Versionnim-aa28e1f-nvfp4
UpdatedSeptember 2, 2026 UTC
Compressed Size181.32 GB

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.