Skip to main content
SearchSearch thousands of GPU-optimized Containers, pretrained Models, SDKs, and Helm charts—ready to accelerate AI, digital twins, and HPC from cloud to edge.
NVIDIA AI Enterprise
NVIDIA AI Enterprise
39
26
9
2
  • NVIDIA NIM
    NVIDIA NIM
    9
  • NIM Container GPUs
    NIM Container GPUs
  • Use Case
    Use Case
    42
    38
    14
    12
    12
    9
    9
    8
    6
    6
    5
    4
    4
    4
    3
    2
    2
    2
    2
    2
    1
    1
    1
    1
    1
  • NVIDIA Platform
    NVIDIA Platform
    30
    27
    19
    15
    10
    8
    8
    7
    5
    5
    4
    4
    3
    3
    3
    2
    2
    2
    1
    1
    1
    1
    1
    1
    1
    1
    1
  • Industry
    Industry
    40
    15
    10
    10
    8
    8
    8
    8
    7
    6
    6
    5
    4
    3
    3
    3
    3
    2
  • Solution
    Solution
    404
    273
    269
    248
    182
    176
    120
    113
    105
    55
    54
    48
    40
    16
    16
    14
    13
    11
    11
    10
    7
    4
    4
    3
    2
    1
    1
    1
  • Publisher
    Publisher
    221
    2
    2
    2
    2
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
    1
  • Policy
    Policy
    5
  • Displaying 273 results
    Docker containers distributed as part of the TAO Toolkit package
    Container
    Triton Inference Server is an open source software that lets teams deploy trained AI models from any framework, from local or cloud storage and on any GPU- or CPU-based infrastructure in the cloud, data center, or embedded devices.
    Container
    PyTorch
    NVIDIA
    PyTorch is a GPU accelerated tensor computational framework. Functionality can be extended with common Python libraries such as NumPy and SciPy. Automatic differentiation is done with a tape-based system at the functional and neural network layer levels.
    Container
    TensorFlow is an open source platform for machine learning. It provides comprehensive tools and libraries in a flexible architecture allowing easy deployment across a variety of platforms and devices.
    Container
    TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs.
    Container
    TensorRT
    NVIDIA
    NVIDIA TensorRT is a C++ library that facilitates high-performance inference on NVIDIA graphics processing units (GPUs). TensorRT takes a trained network and produces a highly optimized runtime engine that performs inference for that network.
    Container
    vLLM
    NVIDIA
    vLLM is a fast and easy-to-use library for LLM inference and serving. The NVIDIA vLLM NGC Container is optimized for GPU acceleration, and contains a validated set of libraries that enable and optimize GPU performance.
    Container
    CUDA is a parallel computing platform and programming model that enhances computing performance using NVIDIA GPUs. CUDA Deep Learning integrates networking and GPU-accelerated libraries like cuDNN, cuTensor, NCCL, HPC-x, and the CUDA Toolkit.
    Container
    NVIDIA Optimized Deep Learning Framework, powered by Apache MXNet is a deep learning framework that allows you to mix the flavors of symbolic programming and imperative programming to maximize efficiency and productivity.
    Container
    Allegro Trains delivers an optimized, seamless, and scalable solution for training on DGX machines with ML-Ops and experiment management features.
    Container
    The Dynamo vLLM runtime image is a containerized build of Dynamo + vLLM which serves as the base runtime environment for vLLM based inference with Dynamo's distributed inference framework.
    Container
    NVIDIA Magnum IO is the I/O technologies from NVIDIA and Mellanox that enable applications at scale. The Magnum IO Developer Environment container allows developers to begin scaling their applications on a laptop, desktop, workstation, or in the cloud.
    Container
    The Dynamo TensorRT-LLM runtime image is a containerized build of Dynamo + TensorRT-LLM which serves as the base runtime environment for tensorrt-llm based inference with Dynamo's distributed inference framework.
    Container
    JAX
    NVIDIA
    JAX is a framework for high-performance numerical computing and machine learning research. It includes Numpy-like APIs, automatic differentiation, XLA acceleration and simple primitives for scaling across GPUs and supports an ecosystem of libraries.
    Container
    CUDA GL
    NVIDIA
    CUDA is a parallel computing platform and programming model that enables dramatic increases in computing performance by harnessing the power of the NVIDIA GPUs. These images extend the CUDA images to include OpenGL support through libglvnd.
    Container
    The Dynamo SGLang runtime image is a containerized build of Dynamo + SGLang which serves as the base runtime environment for sglang based inference with Dynamo's distributed inference framework.
    Container
    A comprehensive Helm chart for deploying the NVIDIA Dynamo operator and its dependencies
    Helm Chart
    TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs.
    Container
    PyG
    NVIDIA
    PyG (PyTorch Geometric) is a library built upon PyTorch to easily write and train Graph Neural Networks (GNNs) for a wide range of applications related to structured data.
    Container
    kubernetes-operator is a container that runs as part of the Dynamo cloud platform. Dynamo cloud is a kubernetes platform for deploying and managing inference services. This container manages the lifecycle of Dynamo inference deployments in kubernetes.
    Container
    Kaldi
    NVIDIA
    Kaldi is an open-source software framework for speech processing.
    Container
    The Variational Autoencoder for collaborative filtering focuses on providing recommendations.
    Resource
    Triton Inference Server PB October 2025 (PB 25h2) offers a 9-month lifecycle for API stability, with monthly patches for high and critical software vulnerabilities.
    Container
    The Merlin PyTorch container allows users to do preprocessing and feature engineering with NVTabular, and then train a deep-learning based recommender system model with PyTorch, and serve the trained model on Triton Inference Server.
    Container
    ...

    NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.