SearchSearch thousands of GPU-optimized Containers, pretrained Models, SDKs, and Helm charts—ready to accelerate AI, digital twins, and HPC from cloud to edge.
NVIDIA AI Enterprise
NVIDIA AI Enterprise
  • NVIDIA NIM
    NVIDIA NIM
  • NIM Container GPUs
    NIM Container GPUs
  • Use Case
    Use Case
    2
  • NVIDIA Platform
    NVIDIA Platform
    10
  • Industry
    Industry
  • Solution
    Solution
    1
    1
    1
  • Publisher
    Publisher
    11
  • Policy
    Policy
  • Displaying 12 results
    End to End workflow for speech to text training with TAO Toolkit and deployment using Riva.
    Resource
    QuartzNet is a Jasper-like network that uses separable convolutions and larger filter sizes. It has comparable accuracy to Jasper while having much fewer parameters. This particular model has 15 blocks each repeated 5 times.
    Model
    Speech To Text (STT) model based on QuartzNet for recognizing Spanish speech.
    Model
    Speech To Text (STT) model based on QuartzNet for recognizing Russian speech.
    Model
    QuartzNet is a Jasper-like network that uses separable convolutions and larger filter sizes. It has comparable accuracy to Jasper while having much fewer parameters. This particular model has 15 blocks each repeated 5 times.
      Model
      Speech To Text (STT) model based on QuartzNet for recognizing French speech.
      Model
      Speech To Text (STT) model based on QuartzNet for recognizing German speech.
      Model
      QuartzNet15x5 model trained on WSJ, LibriSpeech and Mozilla's Common Voice En with NeMo
      Model
      Speech To Text (STT) model based on QuartzNet for recognizing Polish speech.
      Model
      Speech To Text (STT) model based on QuartzNet for recognizing Italian speech.
      Model
      Speech To Text (STT) model based on QuartzNet for recognizing Catalan speech.
      Model
      Collection
      This collection contains NeMo models for Automatic Speech Recognition (ASR): Speech to Text, Speech Classification, Speaker Diarization, Speaker Verification, Speaker Recognition, Command Recognition, Voice Activity Detection