# Google Transformer Paper

The 2017 Google paper 'Attention Is All You Need' introduced the transformer architecture, a neural network design based on multi-head attention that replaced recurrent units, enabling parallel processing and becoming the foundation for modern large language models.

The transformer is a family of artificial neural network architectures based on the multi-head attention mechanism. In this design, input data such as text, images, or audio is converted into a sequence of numerical representations called tokens, and each token is converted into a vector via lookup from a word embedding table. At each layer, each token is contextualized within the scope of the context window with other unmasked tokens via a parallel multi-head attention mechanism, allowing the signal for key tokens to be amplified and less important tokens to be diminished. Because self-attention alone is permutation-invariant, transformers inject positional information, typically through positional encodings or learned positional embeddings, so token order can affect the output.

The original transformer architecture was proposed in the 2017 paper "Attention Is All You Need" by researchers at Google, including Jakob Uszkoreit, Lukasz Kaiser, and Niki Parmar. This paper marked a departure from recurrent neural networks (RNNs) such as long short-term memory (LSTM), which had dominated sequence modeling. Transformers have the advantage of having no recurrent units, therefore requiring less training time than earlier recurrent architectures, and they have since been widely adopted for training large language models (LLMs) on large datasets. Modern transformer designs are commonly grouped into encoder-only, decoder-only, and encoder-decoder variants, depending on whether they are optimized for representation learning, autoregressive generation, or conditional sequence-to-sequence tasks.

## Predecessors and the Problem of Sequential Processing

For many years, sequence modeling and generation was done using plain recurrent neural networks. A well-cited early example was the Elman network (1990). In theory, information from one token can propagate arbitrarily far down the sequence, but in practice the vanishing-gradient problem leaves the model's state at the end of a long sentence without precise, extractable information about preceding tokens. A key breakthrough was LSTM, originally described in a 1995 technical report and formally published in 1997, which introduced gating mechanisms to mitigate the vanishing gradient problem, allowing efficient learning of long-sequence modeling. One key architectural element was the use of multiplicative gating units, in which the outputs of some neurons modulate the outputs of others. These multiplicative units are conceptually distinct from the additive attention mechanism later introduced for sequence-to-sequence models. LSTM became the standard architecture for long sequence modeling until the 2017 publication of transformers. However, LSTM still used sequential processing, operating one token at a time from first to last; they cannot operate in parallel over all tokens in a sequence.

Modern transformers overcome this problem, but unlike RNNs, they require computation time that is quadratic in the size of the context window. The linearly scaling fast weight controller (1992) learns to compute a weight matrix for further processing depending on the input. One of its two networks has "fast weights" or "dynamic links" (1981). A slow neural network learns by gradient descent to generate keys and values for computing the weight changes of the fast neural network which computes answers to queries. This was later shown to be equivalent to the unnormalized linear transformer.

## Attention with Sequence-to-Sequence Models

The idea of encoder-decoder sequence transduction had been developed in the early 2010s; commonly cited as the originators that produced seq2seq are two concurrently published papers from 2014. A 380M-parameter model for machine translation uses two long short-term memories. Its architecture consists of two parts: the encoder is an LSTM that takes in a sequence of tokens and turns it into a vector, and the decoder is another LSTM that converts the vector into a sequence of tokens. Similarly, another 130M-parameter model used gated recurrent units (GRU) instead of LSTM. Later research showed that GRUs are neither better nor worse than LSTMs for seq2seq.

These early seq2seq models had no attention mechanism, and the state vector is accessible only after the last word of the source text was processed. Although in theory such a vector retains the information about the whole original sentence, in practice the information is poorly preserved. This is because the input is processed sequentially by one recurrent network into a fixed-size output vector, which is then processed by another recurrent network into an output. If the input is long, then the output vector would not be able to contain all relevant information, degrading the output. As evidence, reversing the input sentence improved seq2seq translation.

The RNN search model introduced an attention mechanism to seq2seq for machine translation to solve the bottleneck problem of the fixed-size output vector, allowing the model to process long-distance dependencies more easily. The name is because it "emulates searching through a source sentence during decoding a translation". The relative performances were compared between global and local attention model architectures for machine translation, finding that mixed attention had higher quality than global attention, while local attention reduced translation time.

In 2016, Google Translate was revamped to Google Neural Machine Translation, which replaced the previous model based on statistical machine translation. The new model was a seq2seq model where the encoder and the decoder were both 8 layers of bidirectional LSTM. It took nine months to develop, and it outperformed the statistical approach, which took ten years to develop.

## Parallelizing Attention

Seq2seq models with attention, including self-attention, still suffered from the same issue with recurrent networks: they are hard to parallelize, which prevented them from being accelerated on GPUs. In 2016, decomposable attention applied a self-attention mechanism to feedforward networks, which are easy to parallelize, and achieved state-of-the-art results in textual entailment with an order of magnitude fewer parameters than LSTMs. One of its authors, Jakob Uszkoreit, suspected that attention without recurrence would be sufficient for language translation, thus the title "Attention Is All You Need".

The transformer architecture introduced in the 2017 paper replaced recurrent units entirely with multi-head attention, allowing all tokens to be processed in parallel. This parallelization enabled significant speedups in training and led to the development of pre-trained systems such as generative pre-trained transformers (GPTs) and BERT (bidirectional encoder representations from transformers). Transformers have since found applications in large-scale natural language processing, computer vision (vision transformers), reinforcement learning, audio, multimodal learning, robotics, and playing chess. The architecture has become foundational to the field of [artificial intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence), underpinning modern [large language models](https://www.wikiprompt.org/wiki/large-language-model) developed by organizations such as [OpenAI](https://www.wikiprompt.org/wiki/openai), [Anthropic](https://www.wikiprompt.org/wiki/anthropic), and [Google DeepMind](https://www.wikiprompt.org/wiki/google-deepmind).

## Impact and Legacy

The transformer's design has influenced subsequent research and development in [deep learning](https://www.wikiprompt.org/wiki/deep-learning). Its ability to handle long-range dependencies and parallelize training has made it the default architecture for many tasks. The paper's authors, including Ashish Kumar and others, have contributed to various fields, with some moving to other institutions. The architecture has been adapted for diverse domains, from [Waymo](https://www.wikiprompt.org/wiki/waymo)'s autonomous driving to [Groq](https://www.wikiprompt.org/wiki/groq)'s hardware accelerators. The principles of multi-head attention and positional encoding have become standard components in modern neural network design, and the transformer remains a cornerstone of [generative AI](https://www.wikiprompt.org/wiki/generative-ai) systems.

---
Source: https://www.wikiprompt.org/wiki/google-transformer-paper
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-10-07T16:41:38.910196+00:00
