# T5 Paper (2019)

T5 (Text-to-Text Transfer Transformer) is a series of large language models developed by Google AI in 2019, unifying all NLP tasks into a text-to-text format. The models use encoder-decoder Transformer architecture and are pretrained on the Colossal Clean Crawled Corpus (C4).

T5 (Text-to-Text Transfer Transformer) is a series of large language models developed by Google AI and introduced in 2019. The models are based on the [transformer](https://www.wikiprompt.org/wiki/transformer) architecture, specifically an encoder-decoder design where the encoder processes input text and the decoder generates output text. The key innovation of T5 is its unified text-to-text framework, which treats every natural language processing task - from translation to summarization to question answering - as a text-to-text problem, where both inputs and outputs are always text strings.

The original T5 paper, "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer," demonstrated that a single model architecture could achieve state-of-the-art results across multiple NLP benchmarks by pretraining on a massive corpus and then fine-tuning for specific tasks. This approach influenced subsequent developments in [large-language-model](https://www.wikiprompt.org/wiki/large-language-model) research and contributed to the broader field of [generative-ai](https://www.wikiprompt.org/wiki/generative-ai).

## Training

The original T5 models were pretrained on the Colossal Clean Crawled Corpus (C4), a dataset containing text and code scraped from the internet. This pretraining process enabled the models to learn general language understanding and generation abilities. After pretraining, T5 models could be fine-tuned on specific downstream tasks, adapting their knowledge to perform well in various applications.

The pretraining used a denoising objective where the model learned to reconstruct corrupted text. For example, given the input "Thank you <X> me to your party <Y> week", the model was trained to output "<X> for inviting <Y> last <Z>", where <X> and <Y> denote blanks to be filled (called "sentinels" in the original report) and <Z> marks the end of output. The models were also pretrained on tasks such as translation (e.g., "translate English to German: That is good." -> "Das ist gut.") and grammatical acceptability judgments (e.g., "The course is jumping well." -> "not acceptable").

## Architecture

The T5 series encompasses several models with varying sizes, all using the encoder-decoder Transformer architecture. The original paper reported five model variants distinguished by parameter count: T5-small, T5-base, T5-large, T5-3B, and T5-11B. The encoder and decoder always have the same number of layers; for example, T5-small has 6 layers in both encoder and decoder.

Key architectural parameters include the number of layers (n_layer), number of attention heads (n_head), dimension of embedding vectors (d_model), dimension of the feedforward network (d_ff), and dimension of key/value vectors (d_kv). Unlike typical Transformers, the 3B and 11B models do not satisfy the relationship d_model = d_kv × n_head.

Compared to the original Transformer, T5 introduced several minor modifications: layer normalization without additive bias, placing layer normalization outside the residual path, and using relative positional embeddings. All experiments used a WordPiece tokenizer with vocabulary size 32,000, shared across input and output, trained on a mixture of English, German, French, and Romanian data from C4 at a ratio of 10:1:1:1.

## Variants

Several subsequent models used the T5 architecture, with non-standardized naming conventions. Notable variants include:

- T5 1.1 (small, base, large, XL, XXL): Improved versions with roughly equal parameters, using GEGLU activation instead of ReLU, and renamed 3B/11B to XL/XXL with changed shapes.
- LM-adapted T5 (2021): Started from T5 checkpoints and trained further on 100B additional tokens from C4.
- Switch Transformer (2021): A mixture-of-experts variant replacing feedforward layers with mixture-of-expert layers.
- T0 (2021): Started from LM-adapted T5 checkpoints and trained for zero-shot task instruction following.
- ByT5 (2021): A byte-level version operating on UTF-8 bytes without tokenizers, trained on multilingual C4.
- Flan-T5-XL (2022): Instruction-tuned on the FLAN dataset starting from T5 XL.
- T5X (2022): A JAX-based re-implementation of the original TensorFlow codebase.
- UL2 20B (2022): Scaled to 20B parameters with a "mixture of denoisers" objective, trained on C4.
- Flan-UL2 20B (2022): UL2 20B instruction-finetuned on FLAN.
- Pile-T5 (2024): Same architecture but using the Llama tokenizer, trained on The Pile.

## Applications

T5 models have been employed in various applications, including chatbots, machine translation systems, text summarization tools, code generation, and robotics. The encoder-decoder design allows for instruction following: the encoder processes the instruction and the decoder autoregressively generates the reply.

The T5 encoder can also serve as a text encoder, similar to [BERT](https://www.wikiprompt.org/wiki/bert), producing real-number vector sequences for downstream applications. For instance, Google Imagen uses T5-XXL as a text encoder for conditioning a diffusion model, and the AuraFlow diffusion model uses Pile-T5-XL. These applications demonstrate T5's versatility beyond traditional NLP tasks, contributing to advances in [machine-learning](https://www.wikiprompt.org/wiki/machine-learning) and [deep-learning](https://www.wikiprompt.org/wiki/deep-learning).

---
Source: https://www.wikiprompt.org/wiki/t5-paper
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-10-07T16:33:50.223275+00:00
