NVIDIA
STT En Conformer-Transducer XXLarge
Model
NVIDIA
STT En Conformer-Transducer XXLarge

Conformer-Transducer-XXLarge model for English Automatic Speech Recognition, trained on NeMo ASRSET

  • 1 Version
    1.8.0Selected
    04/14/2022 11:56 PM UTC3.44 GB
    Accuracy
    KeyValue
    NSC Part 16.64 %
    Librispeech test-other3.14 %
    Librispeech dev-other3.09 %
    WSJ Eval 921.49 %
    WSJ Dev 932.42 %
    Librispeech dev-clean1.52 %
    Multilingual Librispeech dev (EN)5.29 %
    Librispeech test-clean1.72 %
    Multilingual Librispeech test (EN)5.85 %
    Model
    KeyValue
    Encoder Dimension1024
    Number of Encoder Layers38
    DatasetNeMo ASRSET 2.0
    ARCHITECTUREConformer-Transducer
    Number of Predictor Layers2
    INPUTS16000 KHZ MONO-CHANNEL AUDIO (WAV FILES)
    Number of Weights1B
    OUTPUTSTRANSCRIBED SPEECH

    NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.