SearchSearch thousands of GPU-optimized Containers, pretrained Models, SDKs, and Helm charts—ready to accelerate AI, digital twins, and HPC from cloud to edge.
NVIDIA AI Enterprise
NVIDIA AI Enterprise
NVIDIA NIM
NVIDIA NIM
NIM Container GPUs
NIM Container GPUs
Use Case
Use Case
7
2
NVIDIA Platform
NVIDIA Platform
1
1
1
Industry
Industry
Solution
Solution
3
2
2
1
Publisher
Publisher
10
5
Policy
Policy
Displaying 17 results
Contains files used in rmir creation
Model
WaveGlow model weights pre-trained on the LJ Speech dataset to be used with https://github.com/NVIDIA/waveglow.
Model
End to End workflow for text to speech training with TAO Toolkit and deployment using Riva.
Resource
Mel-Spectrogram prediction conditioned on input text with LJSpeech voice.
Model
GAN-based waveform generator from mel-spectrograms.
Model
NVIDIA Deep Learning Examples
NVIDIA Deep Learning Examples
HiFi-GAN for PyTorch
HiFi-GAN model implements a spectrogram inversion model that allows to synthesize speech waveforms from mel-spectrograms.
Resource
NVIDIA Deep Learning Examples
NVIDIA Deep Learning Examples
Tacotron2 PyTorch checkpoint (AMP)
Tacotron2 PyTorch checkpoint trained with AMP
Model
NVIDIA Deep Learning Examples
NVIDIA Deep Learning Examples
Tacotron2 and Waveglow 2.0 for PyTorch
The Tacotron 2 and WaveGlow model form a text-to-speech system that enables user to synthesise a natural sounding speech from raw transcripts.
Resource
Universal waveform generator from mel-spectrograms.
Model
NVIDIA Deep Learning Examples
NVIDIA Deep Learning Examples
FastPitch 1.0 for PyTorch
The FastPitch model generates mel-spectrograms from raw input text and allows to exert additional control over the synthesized utterances.
Resource
Mel-Spectrogram prediction conditioned on input text with LJSpeech voice.
Model
NVIDIA Deep Learning Examples
NVIDIA Deep Learning Examples
Waveglow PyTorch checkpoint
Waveglow PyTorch checkpoint trained with AMP
Model
FastPitch is a mel-spectrogram generator, designed to be used as the first part of a neural text-to-speech system in conjunction with a neural vocoder
Model
GAN-based waveform generator from mel-spectrograms.
Model
HifiGAN is a neural vocoder model for text-to-speech applications. It is intended as the second part of a two-stage speech synthesis pipeline, with a mel-spectrogram generator such as FastPitch as the first stage.
Model
Collection
This collection contains NeMo models for Text to Speech (TTS)
23
Collection
A collection of easy to use, highly optimized Deep Learning Models for Speech Synthesis. Deep Learning Examples provides Data Scientist and Software Engineers with recipes to Train, fine-tune, and deploy State-of-the-Art Models
210

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.