NVIDIA
NVIDIA
Triton Inference Server PB October 2025 (PB 25h2)
Container
NVIDIA
NVIDIA
Triton Inference Server PB October 2025 (PB 25h2)

Triton Inference Server PB October 2025 (PB 25h2) offers a 9-month lifecycle for API stability, with monthly patches for high and critical software vulnerabilities.

Subscribe to get accessSubscribe to the product below to access this premium content:
NVIDIA AI Enterprise
NVIDIA AI EnterpriseAccelerate your AI agent development
Subscribe Now
Note: You can gain access to hundreds more GPU-optimized artifacts by creating a free NGC account.
Already Subscribed?Log in

What is Triton Inference Server?

Triton Inference Server provides a cloud and edge inferencing solution optimized for both CPUs and GPUs. Triton supports an HTTP/REST and GRPC protocol that allows remote clients to request inferencing for any model being managed by the server. For edge deployments, Triton is available as a shared library with a C API that allows the full functionality of Triton to be included directly in an application. The following Docker images are available:

  1. The 25.08.xx-py3 image contains the Triton Inference Server with support for PyTorch, TensorRT, ONNX and OpenVINO models.

  2. The 25.08.xx-py3-sdk image contains Python and C++ client libraries, client examples, GenAI-Perf, Performance Analyzer and the Model Analyzer.

  3. The 25.08.xx-vllm-python-py3 image contains the Triton Inference Server with support for vLLM and Python backends only.

  4. The 25.08.xx-trtllm-python-py3 image contains the Triton Inference Server with support for TensorRT-LLM and Python backends only.

What Is Triton Inference Server Production Branch October 2025?

The Triton Inference Server Production Branch, exclusively available with NVIDIA AI Enterprise, is a 9-month supported, API-stable branch that includes monthly fixes for high and critical software vulnerabilities. This branch provides a stable and secure environment for building your mission-critical AI applications. The Triton Inference Server production branch releases every six months with a three-month overlap in between two releases.

Getting started with Triton Inference Server Production Branch

Before you start, ensure that your environment is set up by following one of the deployment guides available in the NVIDIA AI Enterprise Documentation.

For an overview of the features included in the Triton Inference Server Production Branch October 2025, please refer to the Release Notes for Triton Inference Server 25.08.

For more information about the Triton Inference Server, see:

Additionally, if you're looking for information on Docker containers and guidance on running a container, review the Containers For Deep Learning Frameworks User Guide.

Government Ready

This ensures the highest level of security for regulated environments, the x86 container image for this branch is:

  • STIG Ubuntu 24.04 hardened
  • Supports FIPS 140-2 / 3 validated crypto / uses libraries that support FIPS crypto

To use this specific hardened image, navigate to the repository's Tags tab and look for the purple label indicating Gov ready displayed alongside the tag. Learn more about NVIDIA's hardened image in the AI Software for Regulated Environments White Paper.

Compatible Infrastructure Software Versions

For the optimized performance, it is highly recommended to deploy the supported NVIDIA AI Enterprise Infrastructure software in conjunction with your AI software. Production Branch - October 2025 (25h2) is compatible with NVIDIA AI Enterprise Infrastructure 7.

OSS License Archive

OSS License Archive contains all project-related licenses. It ensures transparency and compliance with legal requirements, providing detailed information about the terms and conditions associated with the use, modification, and distribution of this project.

Security Vulnerabilities in Open Source Packages

Please review the Security Scanning tab to view the latest security scan results.

For certain open-source vulnerabilities listed in the scan results, NVIDIA provides a response in the form of a Vulnerability Exploitability eXchange (VEX) document. The VEX information can be reviewed and downloaded from the Security Scanning tab.

Get Help

Enterprise Support

Get access to knowledge base articles and support cases or submit a ticket.

NVIDIA AI Enterprise Documentation

Visit the NVIDIA AI Enterprise Documentation Hub for release documentation, deployment guides and more.

NVIDIA Licensing Portal

Go to the NVIDIA Licensing Portal to manage your software licenses. licensing portal for your products. Get Your Licenses

License

By pulling and using the container, you accept the terms and conditions of this End User License Agreement and Product-Specific Terms.

You are responsible for ensuring that your use of NVIDIA provided models complies with all applicable laws.

Governing Terms

The software and materials are governed by the NVIDIA Software and Model Evaluation License.

Publisher
NVIDIA
NVIDIA
Latest Tag25.08.10-vllm-python-py3
UpdatedJune 30, 2026 UTC
Compressed Size8.63 GB
Multinode SupportNo
Multi-Arch SupportYes

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.