# Attention Is All You Need (2017)

A 2017 Google research paper introducing the Transformer architecture, which uses self-attention and became the foundation for most modern large language models.

"Attention Is All You Need" is a 2017 research paper on machine learning authored by eight scientists and engineers working at Google. The paper introduced a new deep learning architecture known as the transformer, based on the attention mechanism proposed in 2014 by Bahdanau et al. The transformer approach it describes has become the main architecture of a wide variety of artificial intelligence systems, including [large language models](https://www.wikiprompt.org/wiki/large-language-model). At the time, the focus of the research was on improving Seq2seq techniques for machine translation, but the authors went further in the paper, foreseeing the technique's potential for other tasks like question answering and what is now known as multimodal [generative AI](https://www.wikiprompt.org/wiki/generative-ai).

Some early examples on which the team tried their Transformer architecture included English-to-German translation, generating Wikipedia articles on "The Transformer", and parsing. These convinced the team that the Transformer is a general-purpose language model, and not just good for translation.

As of 2026, the paper has been cited more than 250,000 times, placing it among the top ten most-cited papers of the 21st century. After the paper was published by Google, each of the eight authors left the company to join other companies or to found startups.

## Background

The authors of the paper are Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Lukasz Kaiser, and Illia Polosukhin. All eight authors were "equal contributors" to the paper; the listed order was randomized (according to the paper itself). After the paper, each of the authors left Google to join other companies or to found startups.

The paper's title is a reference to the song "All You Need Is Love" by the Beatles. The name "Transformer" was picked because Jakob Uszkoreit, one of the paper's authors, liked the sound of that word. An early design document was titled "Transformers: Iterative Self-Attention and Processing for Various Tasks", and included an illustration of six characters from the Transformers franchise. The team was named Team Transformer.

## Methods discussed and introduced

The paper is best known for introducing the Transformer architecture, which underlies most modern large language models (LLMs). A key reason why the architecture is preferred by most modern LLMs is the parallelizability of the architecture over its predecessors. This ensures that the operations necessary for training can be accelerated on a GPU, allowing both faster training times and models of bigger sizes to be trained.

The paper introduced the following mechanisms as part of the development of the transformer architecture.

### Scaled dot-product attention and self-attention

The use of the scaled dot-product attention and self-attention mechanism instead of a recurrent neural network (RNN) or long short-term memory (which rely on recurrence) allows for better performance as described in the following paragraph. The paper described the scaled dot-product attention as follows:

Attention(Query, Key, Value) := softmax(Query × Key^T / √d_k) × Value

where Query, Key, Value are respectively the query, key, value matrices, and d_k is the dimension of the values.

Since the model relies on Query, Key, and Value matrices that come from the same source (i.e., the input sequence or context window), this eliminates the need for RNNs, completely ensuring parallelizability for the architecture. This differs from the original form of the Attention mechanism introduced in 2014. Additionally, the paper discusses the use of an additional scaling factor that was found to be most effective with respect to the dimension of the key vectors (represented as d_k and initially set to 64 within the paper) in the manner shown above.

In the specific context of translation, which the paper focused on, the Query and Key matrices are usually represented in embeddings corresponding to the source language, while the Value matrix corresponds to the target language.

### Multi-head attention

In the self-attention mechanism, queries (Q), keys (K), and values (V) are dynamically generated for each input sequence (typically limited by the size of the context window), allowing the model to focus on different parts of the input sequence at different steps. Multi-head attention enhances this process by introducing multiple parallel attention heads. Each attention head learns different linear projections of the Q, K, and V matrices. This allows the model to capture different aspects of the relationships between words in the sequence simultaneously, rather than focusing on a single aspect.

By doing so, multi-head attention enables the model to attend to information from different representation subspaces at different positions. The paper set the number of heads to 8, with each head having a reduced dimension (d_k = 64), resulting in a total computational cost similar to that of a single full-dimensional attention head.

### Positional encoding

Because the Transformer architecture does not inherently capture the order of tokens in a sequence, the paper introduced positional encodings. These are vectors added to the input embeddings to provide information about the position of each token. The paper used sine and cosine functions of different frequencies, allowing the model to learn to attend by relative position. This was a crucial component for tasks like translation, where word order matters.

### Other components

The paper also incorporated several other techniques that were already known in the deep learning community. It used [layer normalization](https://www.wikiprompt.org/wiki/layer-normalization) to stabilize training, [dropout](https://www.wikiprompt.org/wiki/dropout) for regularization, and the [Adam optimizer](https://www.wikiprompt.org/wiki/adam-optimizer) with a custom learning rate schedule that included a warmup phase followed by decay. The architecture followed an [encoder-decoder](https://www.wikiprompt.org/wiki/encoder-decoder) structure, with stacked layers of multi-head attention and feed-forward networks, and used [residual connections](https://www.wikiprompt.org/wiki/residual-network) around each sub-layer.

## Impact and legacy

The Transformer architecture introduced in the paper quickly became the dominant approach in natural language processing. It replaced recurrent neural networks in many applications, leading to the development of models like BERT and GPT. These models, in turn, underpinned the rise of large language models such as those developed by [OpenAI](https://www.wikiprompt.org/wiki/openai), [Anthropic](https://www.wikiprompt.org/wiki/anthropic), and [Google DeepMind](https://www.wikiprompt.org/wiki/google-deepmind). The architecture's parallelizability made it particularly well-suited for training on specialized hardware like GPUs and TPUs, enabling the scaling of models to unprecedented sizes.

Beyond language, the Transformer has been applied to computer vision, audio processing, and even protein folding, demonstrating its versatility. The paper's influence is reflected in its citation count, which exceeded 250,000 by 2026, making it one of the most cited papers in computer science history.

The authors' subsequent careers also highlight the paper's significance. After leaving Google, they founded or joined various AI startups and research labs, including companies like AI21 Labs, Essential AI, and Inceptive, contributing to the broader AI ecosystem.

## Reception and recognition

The paper was initially presented at the 2017 Conference on Neural Information Processing Systems (NeurIPS) in Long Beach, California, where it received the Best Paper Award. It was praised for its elegant simplification of sequence modeling and its strong empirical results, achieving state-of-the-art performance on machine translation tasks with faster training times than existing recurrent models.

Over time, the paper has been recognized as a foundational work in the field of [deep learning](https://www.wikiprompt.org/wiki/deep-learning). It is often cited as a key milestone in the history of [artificial intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence), alongside other breakthroughs like the [residual network](https://www.wikiprompt.org/wiki/residual-network) and the [Adam optimizer](https://www.wikiprompt.org/wiki/adam-optimizer). The term "attention" itself has become a central concept in AI research, with numerous extensions and variations proposed in subsequent years.

Despite its success, the paper has also faced scrutiny regarding the interpretability of attention mechanisms and the computational costs of large transformers, leading to ongoing research into more efficient architectures and alternative attention mechanisms. Nevertheless, the Transformer remains the backbone of most state-of-the-art AI systems as of 2026.

---
Source: https://www.wikiprompt.org/wiki/attention-is-all-you-need-paper
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-10-07T16:41:52.0414+00:00
