Nvidia NeMo Speech supports all training stages for Nemotron Speech models including Nemotron-ASR, Nemotron-VoiceChat, Parakeet, Canary, MagpieTTS, and more.
What is the NeMo Speech Framework?
NVIDIA NeMo Speech is built for researchers and PyTorch developers working on Speech models including Automatic Speech Recognition (ASR), Text to Speech (TTS), and Speech LLMs. It is designed to help you efficiently create, customize, and deploy new AI models by leveraging existing code and pre-trained model checkpoints.
NeMo Speech supports training and inference for Nemotron-Speech models: Nemotron-ASR, Nemotron-VoiceChat, Parakeet, Canary, MagpieTTS, and more. For accelerated inference and serving, please see our NIMs: https://build.nvidia.com/explore/speech.
Getting Started With NVIDIA NeMo
Refer to the NVIDIA NeMo Speech Documentation page for step-by-step guides to get started quickly with a domain.
Refer to the the Nemotron Speech HuggingFace Collection for quick links to demos and individual checkpoints. For any individual model, please refer to their HuggingFace model page.
Questions? See the current discussions and submit a question.
Report a bug? You can report a bug.
License
Governing Terms: Your use of the NeMo Framework is governed by the NVIDIA Software License Agreement and the Product-Specific Terms for NVIDIA AI Products.