Skip to main content
SearchSearch thousands of GPU-optimized Containers, pretrained Models, SDKs, and Helm charts—ready to accelerate AI, digital twins, and HPC from cloud to edge.
NVIDIA AI Enterprise
NVIDIA AI Enterprise
3
2
1
  • NVIDIA NIM
    NVIDIA NIM
    2
  • NIM Container GPUs
    NIM Container GPUs
  • Use Case
    Use Case
    119
    105
    65
    60
    38
    37
    35
    34
    33
    33
    27
    26
    24
    24
    22
    21
    14
    13
    12
    11
    10
    10
    9
    8
    8
    5
    4
    4
    3
    2
    1
    1
    1
    1
    1
    1
  • NVIDIA Platform
    NVIDIA Platform
    6
    3
    2
    2
    1
    1
    1
    1
    1
    1
    1
    1
  • Industry
    Industry
    4
    4
    2
    2
    1
  • Solution
    Solution
    22
    10
    7
    7
    6
    5
    4
    2
    2
    1
    1
    1
    1
  • Publisher
    Publisher
    29
    1
  • Policy
    Policy
  • Displaying 34 results
    vLLM
    NVIDIA
    vLLM is a fast and easy-to-use library for LLM inference and serving. The NVIDIA vLLM NGC Container is optimized for GPU acceleration, and contains a validated set of libraries that enable and optimize GPU performance.
    Container
    The Llama 3.2 Vision instruction-tuned models are optimized for visual recognition, image reasoning, captioning, and answering general questions about an image.
    Container
    Build a Video Search and Summarization Agent Ingest massive volumes of live or archived videos and extract insights for summarization and interactive Q&A
    Container
    NVIDIA NIM for GPU accelerated Llama 2 7B inference through OpenAI compatible APIs
    Container
    SGLang
    NVIDIA
    SGLang is a fast serving framework for large language models and vision language models. The NVIDIA SGLang NGC Container is optimized for GPU acceleration, and contains a validated set of libraries that enable and optimize GPU performance.
    Container
    The NVIDIA AI-Q Research Assistant Blueprint gives developers a foundational starting point for building a deep research assistant that can run on-premise. The backend container provides the RESTful API service.
    Container
    Base Container for building VSS engine from source
    Container
    The NVIDIA AI-Q Research Assistant Blueprint gives developers a foundational starting point for building a deep research assistant that can run on-premise. The frontend container provides a demo web UI.
    Container
    Helm Chart for NeMo Retriever NVIngest Microservice
    Helm Chart
    Blueprint for the Video Search and Summarization Agent
    Helm Chart
    aiq-agent
    NVIDIA
    NVIDIA AI-Q Intelligence Agent — an enterprise-grade backend agent built on the NVIDIA NeMo Agent Toolkit, providing quick cited answers and in-depth report-style research with modular multi-agent workflows.
    Container
    The NVIDIA AI-Q Research Assistant Blueprint gives developers a foundational starting point for building a deep research assistant that can run on-premise. This container includes a utility to create two default datasets used by the demo web UI.
    Container
    The NVIDIA K8s Developer LLM Operator is an open source and easy to deploy Kubernetes Operator to self-host Generative AI workflows.
    Helm Chart
    AI Inference Service for using VLM (visual language model) on streaming video for greater contextual understanding and natural language interaction
    Container
    NVIDIA AI-Q Blueprint frontend — a modern web UI built with Next.js, React, and NVIDIA KUI Foundations, providing an accessible interface for the AI-Q Blueprint backend with optional OAuth authentication.
    Container
    NeMo Retriever Library is a scalable, performance-oriented document content and metadata extraction microservice.
    Container
    AI-Q Blueprint
    NVIDIA Corporation
    AI-Q Blueprint Helm Chart
    Helm Chart
    Helm chart to deploy the NVIDIA AI-Q Research Assistant Blueprint which gives developers a starting point for building a deep research assistant that can run on-premise, allowing anyone to create detailed research reports using on-premise data.
    Helm Chart
    Jetson Platform Services Reference Workflow & Resources
    Resource
    Question answering model with BERT base encoder finetuned on SQuADv2.0
    Model
    Question answering model with BERT base encoder finetuned on SQuADv1.1
    Model
    Uncased question answering model with Megatron encoder finetuned on SQuADv1.1
    Model
    Question answering model with BERT large encoder finetuned on SQuADv2.0
    Model
    Cased question answering model with Megatron encoder finetuned on SQuADv1.1
    Model

    NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.