SearchSearch thousands of GPU-optimized Containers, pretrained Models, SDKs, and Helm charts—ready to accelerate AI, digital twins, and HPC from cloud to edge.
NVIDIA Enterprise
NVIDIA Enterprise
4
1
1
NVIDIA NIM
NVIDIA NIM
6
NIM Container GPUs
NIM Container GPUs
Use Case
Use Case
119
104
64
61
37
37
34
34
33
33
26
24
23
22
22
21
15
13
11
10
10
8
8
8
6
5
4
4
3
2
1
1
1
1
1
1
NVIDIA Platform
NVIDIA Platform
13
4
4
3
1
1
1
1
1
1
1
Industry
Industry
5
2
2
1
1
Solution
Solution
31
21
14
13
5
5
3
1
1
1
Publisher
Publisher
39
13
2
2
1
1
1
1
Policy
Policy
1
Displaying 64 results
NVIDIA
NVIDIA
PyTorch
PyTorch is a GPU accelerated tensor computational framework. Functionality can be extended with common Python libraries such as NumPy and SciPy. Automatic differentiation is done with a tape-based system at the functional and neural network layer levels.
Container
NVIDIA NeMo™ framework Megatron backend supports pre-training, post-training, and reinforcement learning of LLMs and multi-modal generative AI models with state-of-the-art data processing, model training techniques, and flexible deployment options.
Container
NVIDIA
NVIDIA
vLLM
vLLM is a fast and easy-to-use library for LLM inference and serving. The NVIDIA vLLM NGC Container is optimized for GPU acceleration, and contains a validated set of libraries that enable and optimize GPU performance.
Container
This container houses the Llama-3.3-Nemotron-Super-49B-v1.5, which is a significantly upgraded version of Llama-3.3-Nemotron-Super-49B-v1 and is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct
Container
Scripts and utilities for getting started with Riva Speech Skills
Resource
NVIDIA
NVIDIA
Kaldi
Kaldi is an open-source software framework for speech processing.
Container
NVIDIA AI Enterprise
This container houses the **Llama-3.3-Nemotron-Super-49B-v1.5 PB 25h2**, which is a significantly upgraded version of Llama-3.3-Nemotron-Super-49B-v1 and is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct
Container
NVIDIA
NVIDIA
SGLang
SGLang is a fast serving framework for large language models and vision language models. The NVIDIA SGLang NGC Container is optimized for GPU acceleration, and contains a validated set of libraries that enable and optimize GPU performance.
Container
Nemotron Content Safety Reasoning 4B is a Large Language Model (LLM) classifier designed to function as a dynamic and adaptable guardrail for content safety and dialogue moderation (topic-following).
Container
This container houses the model MiMo-V2-Flash.
Container
The GPT-OSS-120b-Turbo NIM container packages OpenAI's GPT-OSS-120b large language model, a sparse Mixture of Experts (MoE) architecture with 120B total parameters and 5.1B active parameters, as an NVIDIA NIM microservice.
Container
Nemotron-3-Ultra-550B-A55B NIM container packages NVIDIA's large language model featuring a hybrid Latent Mixture-of-Experts (LatentMoE) architecture with Multi-Token Prediction (MTP) layers.
Container
Jupyter Notebook example for Question Answering with BERT for TensorFlow
Resource
Jupyter Notebooks for BERT Pre-training, Fine-Tuning and Inference profiling and optimization via TensorFlow, AMP, XLA, DLProf, TF-TRT and Triton.
Container
Base environment used in the NVIDIA NeMo projects of the NVIDIA Deep Learning Institute (DLI) course, "Building Transformer-Based Natural Language Processing Applications". This container also includes a "Next Steps" project.
Container
NVIDIA Corporation
NVIDIA Corporation
AI-Q Blueprint
AI-Q Blueprint Helm Chart
Helm Chart
Base environment of the NVIDIA Deep Learning Institute (DLI) course, "Building Conversational AI Applications". This container also includes a "Next Steps" project.
Container
NVIDIA NeMo Microservices
NVIDIA
NVIDIA
deplot
NVIDIA NIM for GPU accelerated Deplot inference through OpenAI compatible APIs
Container
University of Florida Health
GatorTron-OG
GatorTron-OG is a Megatron BERT model trained on pre-trained on de-identified clinical notes from the University of Florida Health System.
Model
Fine-tune a pre-trained BERT model with the SQuAD dataset, optimize for inference using TensorRT and deploy with Triton Inference Server on Google Cloud AI Platform using Custom Containers
Resource
Megatron pretrained on uncased biomedical dataset PubMed with 345 million parameters.
Model
End to End workflow for question answering starting with training in TAO Toolkit and deployment using Riva.
Resource
NVIDIA's BERT leverages mixed precision arithmetic and Tensor Cores on A100, V100 and T4 GPUs for faster training while maintaining target accuracy. This notebook demonstrates BERT Question Answering Fine-Tuning with Mixed Precision on SQuaD 2.0 Dataset.
Resource
End to End sample workflow for Text Classification starting with training in TAO Toolkit and deployment using Riva.
Resource

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.