NVIDIA
STT Fa FastConformer Hybrid Transducer-CTC Large
Model
NVIDIA
STT Fa FastConformer Hybrid Transducer-CTC Large

This collection contains the large version (114M) of the Persian speech recognition model with a FastConformer encoder and a Hybrid decoder (joint RNNT-CTC loss). The model has a vocab size of 1024.

  • 1 Version
    1.21.0Selected
    11/07/2023 6:33 PM UTC437.96 MB
    Accuracy
    KeyValue
    CTC WER CV-Fa dev (custom)13.18
    RNNT CER CV-Fa dev (custom)3.89
    CTC CER CV-Fa test (custom)3.85
    RNNT WER CV-Fa test (custom)15.48
    CTC WER CV-Fa test (custom)13.16
    RNNT WER CV-Fa dev (custom)15.44
    RNNT CER CV-Fa test (custom)4.63
    CTC CER CV-Fa dev (custom)3.38
    Model
    KeyValue
    ArchitectureFastConformer Hybrid Transducer-CTC Large
    OutputsSequence of BPE tokens representing recognized text
    InputsSequence of 16kHz single-channel audio samples

    NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.