Skip to main content
NVIDIA
Riva TTS A²-Flow for Nv IGI SDK
Model
NVIDIA
Riva TTS A²-Flow for Nv IGI SDK

Riva TTS A²-Flow model for the NVIDIA In-Game Inferencing (NVIGI) SDK.

This model is backed by NVIDIA's Plus Plus (++) Promise
to learn more about the quality of the datasets used to train this model.
FieldResponse
Intended Application & Domain:Speech Synthesis
Model TaskSpeech Synthesis and Voice Characterization
Intended UsersThis model is intended for developers building interactive call centers, virtual assistants, and language learning assistants to improve pronunciation, automatically generate voice-overs, narrate or comment on videos, and provide audio alternatives for visually impaired users or people with light sensitivity.
Model OutputAudio of shape (batch x time) in wav format
Describe how the model worksModel takes input text and outputs an audio representation of the text. It can be used as a zero-shot voice characterization model. When given a reference audio sample to replicate along with an input text, the produced synthetic audio will be similar to this reference.
Name the adversely impacted groups this has been tested to deliver comparable outcomes regardless ofGender, Age (including people in older age brackets)
Technical LimitationsModel only has the capacity to produce a voice in the languages, dialects and gender(s) in which it is trained. This model makes no effort to moderate or modify input text. Languages that are underrepresented may not sound as natural.
Verified to have met prescribed NVIDIA quality standardsYes
Performance Metrics% preference when compared with available alternatives
word error rate (wer)
character error rate (cer)
mean opinion score (MOS)
Potential Known RisksThis model has the ability to replicate the characteristics of an individual's voice but may unnaturally synthesize vocabulary not included in the pronunciation dictionary or omit phonetic symbols not used in training.
Licensing:https://docs.nvidia.com/ai-foundation-models-community-license.pdf