The GNMT v2 model is an improved version of the first Google's Neural Machine Translation System with a modified attention mechanism.
Performance
Results
Training Accuracy Results
Results were obtained by running the train.py script with the default
batch size = 128 per GPU in the pytorch-19.01-py3 Docker container.
NVIDIA DGX-1 (8x Tesla V100 16G)
Command used to launch the training:
python3 -m launch train.py --seed 2 --train-global-batch-size 1024
| number of GPUs | batch size/GPU | mixed precision BLEU | fp32 BLEU | mixed precision training time | fp32 training time |
|---|---|---|---|---|---|
| 1 | 128 | 24.59 | 24.71 | 264.4 minutes | 824.4 minutes |
| 4 | 128 | 24.30 | 24.45 | 89.5 minutes | 230.8 minutes |
| 8 | 128 | 24.45 | 24.48 | 46.2 minutes | 116.6 minutes |
NVIDIA DGX-2 (16x Tesla V100 32G)
Commands used to launch the training:
for 1,4,8 GPUs:
python3 -m launch train.py --seed 2 --train-global-batch-size 1024
for 16 GPUs:
python3 -m launch train.py --seed 2 --train-global-batch-size 2048
| number of GPUs | batch size/GPU | mixed precision BLEU | fp32 BLEU | mixed precision training time | fp32 training time |
|---|---|---|---|---|---|
| 1 | 128 | 24.59 | 24.71 | 265.0 minutes | 825.1 minutes |
| 4 | 128 | 24.69 | 24.33 | 87.4 minutes | 216.3 minutes |
| 8 | 128 | 24.50 | 24.47 | 49.6 minutes | 113.5 minutes |
| 16 | 128 | 24.22 | 24.16 | 26.3 minutes | 58.6 minutes |

Training Stability Test
The GNMT v2 model was trained for 6 epochs, starting from 50 different initial random seeds. After each training epoch the model was evaluated on the test dataset and the BLEU score was recorded. The training was performed in the pytorch-19.01-py3 Docker container on NVIDIA DGX-1 with 8 Tesla V100 16G GPUs. The following table summarizes results of the stability test.

BLEU scores after each training epoch for different initial random seeds
| epoch | average | stdev | minimum | maximum | median |
|---|---|---|---|---|---|
| 1 | 19.954 | 0.326 | 18.710 | 20.490 | 20.020 |
| 2 | 21.734 | 0.222 | 21.220 | 22.120 | 21.765 |
| 3 | 22.502 | 0.223 | 21.960 | 22.970 | 22.485 |
| 4 | 23.004 | 0.221 | 22.350 | 23.430 | 23.020 |
| 5 | 24.201 | 0.146 | 23.900 | 24.480 | 24.215 |
| 6 | 24.423 | 0.159 | 24.070 | 24.820 | 24.395 |
Training Performance Results
All results were obtained by running the train.py training script in the
pytorch-19.01-py3 Docker container. Performance numbers (in tokens per second)
were averaged over an entire training epoch.
NVIDIA DGX-1 (8x Tesla V100 16G)
| number of GPUs | batch size/GPU | mixed precision tokens/s | fp32 tokens/s | mixed precision speedup | mixed precision multi-gpu strong scaling | fp32 multi-gpu strong scaling |
|---|---|---|---|---|---|---|
| 1 | 128 | 66050 | 21346 | 3.094 | 1.000 | 1.000 |
| 4 | 128 | 196174 | 76083 | 2.578 | 2.970 | 3.564 |
| 8 | 128 | 387282 | 153697 | 2.520 | 5.863 | 7.200 |
NVIDIA DGX-2 (16x Tesla V100 32G)
| number of GPUs | batch size/GPU | mixed precision tokens/s | fp32 tokens/s | mixed precision speedup | mixed precision multi-gpu strong scaling | fp32 multi-gpu strong scaling |
|---|---|---|---|---|---|---|
| 1 | 128 | 65830 | 22695 | 2.901 | 1.000 | 1.000 |
| 4 | 128 | 200886 | 81224 | 2.473 | 3.052 | 3.579 |
| 8 | 128 | 362612 | 156536 | 2.316 | 5.508 | 6.897 |
| 16 | 128 | 738521 | 314831 | 2.346 | 11.219 | 13.872 |
Inference Performance Results
All results were obtained by running the translate.py script in the
pytorch-19.01-py3 Docker container on NVIDIA DGX-1. Inference benchmark was run
on a single Tesla V100 16G GPU. The benchmark requires a checkpoint from a fully
trained model.
Command to launch the inference benchmark:
python3 translate.py --input data/wmt16_de_en/newstest2014.tok.bpe.32000.en \
--reference data/wmt16_de_en/newstest2014.de --output /tmp/output \
--model results/gnmt/model_best.pth --batch-size 32 128 512 \
--beam-size 1 2 5 10 --math fp16 fp32
| batch size | beam size | mixed precision BLEU | fp32 BLEU | mixed precision tokens/s | fp32 tokens/s |
|---|---|---|---|---|---|
| 32 | 1 | 23.18 | 23.18 | 23571 | 19462 |
| 32 | 2 | 24.09 | 24.12 | 15303 | 12345 |
| 32 | 5 | 24.63 | 24.62 | 13644 | 7725 |
| 32 | 10 | 24.50 | 24.48 | 11049 | 5359 |
| 128 | 1 | 23.17 | 23.18 | 73429 | 42272 |
| 128 | 2 | 24.07 | 24.12 | 43373 | 23131 |
| 128 | 5 | 24.69 | 24.63 | 29646 | 12525 |
| 128 | 10 | 24.45 | 24.48 | 19100 | 6886 |
| 512 | 1 | 23.17 | 23.18 | 135333 | 48962 |
| 512 | 2 | 24.08 | 24.12 | 74367 | 27308 |
| 512 | 5 | 24.60 | 24.63 | 39217 | 12674 |
| 512 | 10 | 24.54 | 24.48 | 21433 | 6640 |