SearchSearch thousands of GPU-optimized Containers, pretrained Models, SDKs, and Helm charts—ready to accelerate AI, digital twins, and HPC from cloud to edge.
NVIDIA Enterprise
NVIDIA Enterprise
3
2
2
1
NVIDIA NIM
NVIDIA NIM
2
NIM Container GPUs
NIM Container GPUs
Use Case
Use Case
119
104
64
61
37
37
34
34
33
32
26
24
23
22
22
21
15
13
11
10
10
8
8
8
6
4
4
4
3
2
1
1
1
1
1
1
NVIDIA Platform
NVIDIA Platform
6
3
2
2
1
1
1
1
1
1
1
1
Industry
Industry
4
4
2
2
1
Solution
Solution
22
10
7
7
6
5
4
2
2
1
1
1
1
Publisher
Publisher
29
1
Policy
Policy
Displaying 34 results
NVIDIA
NVIDIA
vLLM
vLLM is a fast and easy-to-use library for LLM inference and serving. The NVIDIA vLLM NGC Container is optimized for GPU acceleration, and contains a validated set of libraries that enable and optimize GPU performance.
Container
NVIDIA Developer Program
The Llama 3.2 Vision instruction-tuned models are optimized for visual recognition, image reasoning, captioning, and answering general questions about an image.
Container
Build a Video Search and Summarization Agent Ingest massive volumes of live or archived videos and extract insights for summarization and interactive Q&A
Container
NVIDIA Developer Program
NVIDIA NIM for GPU accelerated Llama 2 7B inference through OpenAI compatible APIs
Container
NVIDIA
NVIDIA
SGLang
SGLang is a fast serving framework for large language models and vision language models. The NVIDIA SGLang NGC Container is optimized for GPU acceleration, and contains a validated set of libraries that enable and optimize GPU performance.
Container
The NVIDIA AI-Q Research Assistant Blueprint gives developers a foundational starting point for building a deep research assistant that can run on-premise. The backend container provides the RESTful API service.
Container
Base Container for building VSS engine from source
Container
NVIDIA NeMo Microservices
Helm Chart for NeMo Retriever NVIngest Microservice
Helm Chart
The NVIDIA AI-Q Research Assistant Blueprint gives developers a foundational starting point for building a deep research assistant that can run on-premise. The frontend container provides a demo web UI.
Container
Blueprint for the Video Search and Summarization Agent
Helm Chart
The NVIDIA K8s Developer LLM Operator is an open source and easy to deploy Kubernetes Operator to self-host Generative AI workflows.
Helm Chart
The NVIDIA AI-Q Research Assistant Blueprint gives developers a foundational starting point for building a deep research assistant that can run on-premise. This container includes a utility to create two default datasets used by the demo web UI.
Container
AI Inference Service for using VLM (visual language model) on streaming video for greater contextual understanding and natural language interaction
Container
NVIDIA
NVIDIA
aiq-agent
NVIDIA AI-Q Intelligence Agent — an enterprise-grade backend agent built on the NVIDIA NeMo Agent Toolkit, providing quick cited answers and in-depth report-style research with modular multi-agent workflows.
Container
NVIDIA AI-Q Blueprint frontend — a modern web UI built with Next.js, React, and NVIDIA KUI Foundations, providing an accessible interface for the AI-Q Blueprint backend with optional OAuth authentication.
Container
Helm chart to deploy the NVIDIA AI-Q Research Assistant Blueprint which gives developers a starting point for building a deep research assistant that can run on-premise, allowing anyone to create detailed research reports using on-premise data.
Helm Chart
Jetson Platform Services Reference Workflow & Resources
Resource
NVIDIA Corporation
NVIDIA Corporation
AI-Q Blueprint
AI-Q Blueprint Helm Chart
Helm Chart
NeMo Retriever Library is a scalable, performance-oriented document content and metadata extraction microservice.
Container
Question answering model with BERT base encoder finetuned on SQuADv2.0
Model
Question answering model with BERT base encoder finetuned on SQuADv1.1
Model
Uncased question answering model with Megatron encoder finetuned on SQuADv1.1
Model
Question answering model with BERT large encoder finetuned on SQuADv2.0
Model
Cased question answering model with Megatron encoder finetuned on SQuADv1.1
Model

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.