SearchSearch thousands of GPU-optimized Containers, pretrained Models, SDKs, and Helm charts—ready to accelerate AI, digital twins, and HPC from cloud to edge.
NVIDIA AI Enterprise
NVIDIA AI Enterprise
13
13
7
  • NVIDIA NIM
    NVIDIA NIM
    13
  • NIM Container GPUs
    NIM Container GPUs
    3
    1
    1
  • Use Case
    Use Case
    119
    104
    64
    61
    37
    37
    34
    34
    33
    33
    26
    25
    24
    23
    22
    22
    15
    13
    11
    10
    10
    8
    8
    8
    8
    5
    4
    4
    3
    2
    1
    1
    1
    1
    1
    1
  • NVIDIA Platform
    NVIDIA Platform
    33
    15
    9
    8
    6
    2
    1
    1
    1
    1
  • Industry
    Industry
    2
    1
    1
    1
    1
  • Solution
    Solution
    65
    42
    37
    7
    4
    3
    1
    1
    1
  • Publisher
    Publisher
    101
    2
  • Policy
    Policy
  • Displaying 119 results
    Triton Inference Server is an open source software that lets teams deploy trained AI models from any framework, from local or cloud storage and on any GPU- or CPU-based infrastructure in the cloud, data center, or embedded devices.
    Container
    Riva Speech Skills is a scalable Conversational AI service platform.
    Container
    NVIDIA
    NVIDIA
    Kaldi
    Kaldi is an open-source software framework for speech processing.
    Container
    NVIDIA Developer Program
    Robust Speech Recognition via Large-Scale Weak Supervision.
    Container
    Scripts and utilities for getting started with Riva Speech Skills
    Resource
    NVIDIA Developer Program
    Accurate and optimized English transcriptions with punctuation and word timestamps
    Container
    NVIDIA Developer Program
    Accurate and optimized Vietnamese English transcriptions with punctuation and word timestamps
    Container
    NVIDIA Developer Program
    Accurate and optimized Mandarin English transcriptions with punctuation and word timestamps
    Container
    NVIDIA Developer Program
    Nemotron ASR Streaming
    Container
    NVIDIA Developer Program
    Parakeet-0.6B hybrid-RNNT training with CTC head for zh-tw streaming ASR. Supports Taiwanese Mandarin (zh-tw) and light code-switching capability for Taiwanese Mandarin-English.
    Container
    Multi-scale Diarization Decoder (MSDD) model for speaker diarization of telephone conversations
    Model
    The Domain Specific - NeMo Automatic Speech Recognition (ASR) Application facilitates training, evaluation and performance comparison of ASR models. This NeMo application enables you to train or fine-tune pre-trained ASR models with your own data.
    Container
    ASR + BERT QA interactive chatbot demo for Jetson
    Container
    Citrinet 1024 with kernel scaling factor (gamma) of 25% trained on ASR Set dataset
    Model
    End to End workflow for speech to text training with TAO Toolkit and deployment using Riva.
    Resource
    Citrinet-1024 model with kernel scaling factor (gamma) of 25%, which has been trained on the open-source Aishell-2 Mandarin Chinese corpus.
    Model
    End to End workflow for speech to text citrinet training with TAO Toolkit and deployment using Riva.
    Resource
    This model card contains a Small Audio Codec model trained on the Libri-Light audiobook recordings dataset, comprising approximately 60,000 hours of English language speech with a 16kHz sampling rate.
    Model
    Conformer-CTC-XLarge model for English Automatic Speech Recognition, Trained on NeMo ASRSET
    Model
    German Citrinet 1024 model.
    Model
    Conformer-CTC-Large model for English Automatic Speech Recognition, Trained with NeMo on LibriSpeech dataset
    Model
    AmberNet Lang ID model for Spoken Language Identification
    Model
    Conformer-CTC-Large model for Russian Automatic Speech Recognition, trained on Mozilla Common Voice 10.0 (Russian), Golos (Russian), Russian LibriSpeech (RuLS) and SOVA (RuAudiobooksDevices, RuDevices) datasets.
    Model
    The large version (114M) of the Multilingual speech recognition model with a FastConformer encoder and a Hybrid decoder (joint RNNT-CTC loss). The model has a vocab size of 2560 and emits text with punctuation and capitalization.
    Model

    NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.