NVIDIA
Transformer for PyTorch
Resource
NVIDIA
Transformer for PyTorch

This implementation of Transformer model architecture is based on the optimized implementation in Fairseq NLP toolkit.

Changelog

January 2019

  • initial commit, forked from fairseq

May 2019:

  • add mid-training SacreBLEU evaluation. Better handling of OOMs.

June 2019

  • new README

July 2019

  • replace custom fused operators with jit functions

August 2019

  • add basic AMP support

Known issues

  • Course of a training heavily depends on a random seed. There is high variance in the time required to reach a certain BLEU score. Also the highest BLEU score value observed vary between runs with different seeds.
  • Translations produced by training script during online evaluation may differ from those produced by generate.py script. It is probably a format conversion issue.

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.