NVIDIA
NVIDIA
STT Fa FastConformer Hybrid Transducer-CTC Large
Model
NVIDIA
NVIDIA
STT Fa FastConformer Hybrid Transducer-CTC Large

This collection contains the large version (114M) of the Persian speech recognition model with a FastConformer encoder and a Hybrid decoder (joint RNNT-CTC loss). The model has a vocab size of 1024.

1 Version
1.21.0Selected
11/07/2023 6:33 PM UTC437.96 MB
Accuracy
KeyValue
CTC WER CV-Fa dev (custom)13.18
RNNT CER CV-Fa dev (custom)3.89
CTC CER CV-Fa test (custom)3.85
RNNT WER CV-Fa test (custom)15.48
CTC WER CV-Fa test (custom)13.16
RNNT WER CV-Fa dev (custom)15.44
RNNT CER CV-Fa test (custom)4.63
CTC CER CV-Fa dev (custom)3.38
Model
KeyValue
ArchitectureFastConformer Hybrid Transducer-CTC Large
OutputsSequence of BPE tokens representing recognized text
InputsSequence of 16kHz single-channel audio samples