Google researchers publish the Transformer paper
The paper introduced an encoder–decoder architecture built entirely on attention, replacing recurrent computation with parallelizable self-attention.
What happened
Event details
On June 12, 2017, Google researchers submitted “Attention Is All You Need.” The paper introduced the Transformer, whose encoder and decoder stack self-attention and feed-forward layers instead of recurrent or convolutional networks. The original experiments focused on machine translation. The big configuration scored 28.4 BLEU on WMT 2014 English–German and 41.8 on English–French. Google later made a training implementation available through Tensor2Tensor, but did not release “Transformer” as a single general-purpose pretrained checkpoint.
Assessment
Why it matters
The Transformer changed the basic structure of sequence modeling and became a core architecture across language, vision, audio, and multimodal systems.
Availability
Access notes
The paper and Tensor2Tensor training implementation are public. The original release was not a single official package of general-purpose pretrained weights.