Google researchers publish “Attention Is All You Need”
Eight researchers at Google posted a paper to arXiv describing the transformer, a network built on attention alone, with no recurrence. It reported 28.4 BLEU on the WMT 2014 English-to-German task, trained on eight GPUs in three and a half days.
Why it mattered Almost every large language model released since is a transformer or a close relative of one. What the paper changed was the architecture rather than any single model built on it.
Machine translation in 2017 was built on recurrent networks, which read a sentence one word at a time and carry a summary of what came before. Reading in order is a hard constraint: the work cannot be spread across many processors, and information from the start of a long sentence tends to fade before the end of it.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser and Illia Polosukhin proposed removing recurrence altogether. Their model used only attention, a mechanism that lets every word in a sentence weigh every other word directly, however far apart they sit. Because those comparisons are independent of each other, they run at the same time, and the whole sentence is processed at once.
The results were reported on translation. The larger model reached 28.4 BLEU on the WMT 2014 English-to-German task and 41.8 on English-to-French, and it was trained on eight GPUs in three and a half days, which the authors put at a small fraction of the training cost of the best results published before it.
The architecture spread first through translation and summarization, then through everything else. BERT in 2018 and the GPT series from 2018 onward were transformers; so is nearly every system released since that writes, codes, or answers questions, including ChatGPT five years later. Attention scales with the square of the sequence length, which has made context length one of the standing problems of the field ever since.
The paper was posted to arXiv on 12 June 2017 and presented at the NIPS conference in Long Beach that December. Google Research described the work on its own blog in August, alongside the release of the Tensor2Tensor library.