Skip to main content
NVIDIA
GNMT v2 for PyTorch
Resource
NVIDIA
GNMT v2 for PyTorch

The GNMT v2 model is an improved version of the first Google's Neural Machine Translation System with a modified attention mechanism.

Performance

Results

Training Accuracy Results

Results were obtained by running the train.py script with the default batch size = 128 per GPU in the pytorch-19.01-py3 Docker container.

NVIDIA DGX-1 (8x Tesla V100 16G)

Command used to launch the training:

python3 -m launch train.py --seed 2 --train-global-batch-size 1024
number of GPUsbatch size/GPUmixed precision BLEUfp32 BLEUmixed precision training timefp32 training time
112824.5924.71264.4 minutes824.4 minutes
412824.3024.4589.5 minutes230.8 minutes
812824.4524.4846.2 minutes116.6 minutes
NVIDIA DGX-2 (16x Tesla V100 32G)

Commands used to launch the training:

for 1,4,8 GPUs:
python3 -m launch train.py --seed 2 --train-global-batch-size 1024
for 16 GPUs:
python3 -m launch train.py --seed 2 --train-global-batch-size 2048
number of GPUsbatch size/GPUmixed precision BLEUfp32 BLEUmixed precision training timefp32 training time
112824.5924.71265.0 minutes825.1 minutes
412824.6924.3387.4 minutes216.3 minutes
812824.5024.4749.6 minutes113.5 minutes
1612824.2224.1626.3 minutes58.6 minutes

TrainingLoss

Training Stability Test

The GNMT v2 model was trained for 6 epochs, starting from 50 different initial random seeds. After each training epoch the model was evaluated on the test dataset and the BLEU score was recorded. The training was performed in the pytorch-19.01-py3 Docker container on NVIDIA DGX-1 with 8 Tesla V100 16G GPUs. The following table summarizes results of the stability test.

TrainingAccuracy

BLEU scores after each training epoch for different initial random seeds
epochaveragestdevminimummaximummedian
119.9540.32618.71020.49020.020
221.7340.22221.22022.12021.765
322.5020.22321.96022.97022.485
423.0040.22122.35023.43023.020
524.2010.14623.90024.48024.215
624.4230.15924.07024.82024.395

Training Performance Results

All results were obtained by running the train.py training script in the pytorch-19.01-py3 Docker container. Performance numbers (in tokens per second) were averaged over an entire training epoch.

NVIDIA DGX-1 (8x Tesla V100 16G)
number of GPUsbatch size/GPUmixed precision tokens/sfp32 tokens/smixed precision speedupmixed precision multi-gpu strong scalingfp32 multi-gpu strong scaling
112866050213463.0941.0001.000
4128196174760832.5782.9703.564
81283872821536972.5205.8637.200
NVIDIA DGX-2 (16x Tesla V100 32G)
number of GPUsbatch size/GPUmixed precision tokens/sfp32 tokens/smixed precision speedupmixed precision multi-gpu strong scalingfp32 multi-gpu strong scaling
112865830226952.9011.0001.000
4128200886812242.4733.0523.579
81283626121565362.3165.5086.897
161287385213148312.34611.21913.872

Inference Performance Results

All results were obtained by running the translate.py script in the pytorch-19.01-py3 Docker container on NVIDIA DGX-1. Inference benchmark was run on a single Tesla V100 16G GPU. The benchmark requires a checkpoint from a fully trained model.

Command to launch the inference benchmark:

python3 translate.py --input data/wmt16_de_en/newstest2014.tok.bpe.32000.en \
  --reference data/wmt16_de_en/newstest2014.de --output /tmp/output \
  --model results/gnmt/model_best.pth --batch-size 32 128 512 \
  --beam-size 1 2 5 10 --math fp16 fp32
batch sizebeam sizemixed precision BLEUfp32 BLEUmixed precision tokens/sfp32 tokens/s
32123.1823.182357119462
32224.0924.121530312345
32524.6324.62136447725
321024.5024.48110495359
128123.1723.187342942272
128224.0724.124337323131
128524.6924.632964612525
1281024.4524.48191006886
512123.1723.1813533348962
512224.0824.127436727308
512524.6024.633921712674
5121024.5424.48214336640

Modal Content

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.