Model
This model performs joint intent classification and slot filling, directly from audio input. The model treats the problem as an audio-to-text problem, where the output text is the flattened string representation of the semantics annotation.
Use the NGC CLI to download:
Copied!
1 Version
1.13.0Selected10/06/2022 10:56 PM UTC489.12 MB Copied!
1.13.0Selected
10/06/2022 10:56 PM UTC489.12 MB
Copied!
Accuracy
| Key | Value |
|---|---|
| SLURP-Metrics Precsion (Test) | 84.31 |
| Entity F-1 (Test) | 76.89 |
| SLURP-Metrics Recall (Test) | 80.33 |
| Entity Precsion (Test) | 78.95 |
| Entity Recall (Test) | 74.93 |
| SLURP-Metrics F-1 (Test) | 82.27 |
| Intent Accuracy (Test) | 90.14 |
Model
| Key | Value |
|---|---|
| Encoder Dimension | 512 |
| Number of Layers | 20 |
| Architecture | Conformer-Transformer-Large |
| Dataset | SLURP |
| Outputs | Semantics in English text |
| Inputs | Speech in English |
| Number of Weights | 127M |