T5 Paper (2019)

T5 (Text-to-Text Transfer Transformer) is a series of large language models developed by Google AI in 2019, unifying all NLP tasks into a text-to-text format. The models use encoder-decoder Transformer architecture and are pretrained on the Colossal Clean Crawled Corpus (C4).

T5 (Text-to-Text Transfer Transformer) is a series of large language models developed by Google AI and introduced in 2019. The models are based on the Transformer (architecture) architecture, specifically an encoder-decoder design where the encoder processes input text and the decoder generates output text. The key innovation of T5 is its unified text-to-text framework, which treats every natural language processing task - from translation to summarization to question answering - as a text-to-text problem, where both inputs and outputs are always text strings.

The original T5 paper, "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer," demonstrated that a single model architecture could achieve state-of-the-art results across multiple NLP benchmarks by pretraining on a massive corpus and then fine-tuning for specific tasks. This approach influenced subsequent developments in Large language model research and contributed to the broader field of Generative AI.

Training

The original T5 models were pretrained on the Colossal Clean Crawled Corpus (C4), a dataset containing text and code scraped from the internet. This pretraining process enabled the models to learn general language understanding and generation abilities. After pretraining, T5 models could be fine-tuned on specific downstream tasks, adapting their knowledge to perform well in various applications.

The pretraining used a denoising objective where the model learned to reconstruct corrupted text. For example, given the input "Thank you <X> me to your party <Y> week", the model was trained to output "<X> for inviting <Y> last <Z>", where <X> and <Y> denote blanks to be filled (called "sentinels" in the original report) and <Z> marks the end of output. The models were also pretrained on tasks such as translation (e.g., "translate English to German: That is good." -> "Das ist gut.") and grammatical acceptability judgments (e.g., "The course is jumping well." -> "not acceptable").

Architecture

The T5 series encompasses several models with varying sizes, all using the encoder-decoder Transformer architecture. The original paper reported five model variants distinguished by parameter count: T5-small, T5-base, T5-large, T5-3B, and T5-11B. The encoder and decoder always have the same number of layers; for example, T5-small has 6 layers in both encoder and decoder.

Key architectural parameters include the number of layers (n_layer), number of attention heads (n_head), dimension of embedding vectors (d_model), dimension of the feedforward network (d_ff), and dimension of key/value vectors (d_kv). Unlike typical Transformers, the 3B and 11B models do not satisfy the relationship d_model = d_kv × n_head.

Compared to the original Transformer, T5 introduced several minor modifications: layer normalization without additive bias, placing layer normalization outside the residual path, and using relative positional embeddings. All experiments used a WordPiece tokenizer with vocabulary size 32,000, shared across input and output, trained on a mixture of English, German, French, and Romanian data from C4 at a ratio of 10:1:1:1.

Variants

Several subsequent models used the T5 architecture, with non-standardized naming conventions. Notable variants include:

  • T5 1.1 (small, base, large, XL, XXL): Improved versions with roughly equal parameters, using GEGLU activation instead of ReLU, and renamed 3B/11B to XL/XXL with changed shapes.
  • LM-adapted T5 (2021): Started from T5 checkpoints and trained further on 100B additional tokens from C4.
  • Switch Transformer (2021): A mixture-of-experts variant replacing feedforward layers with mixture-of-expert layers.
  • T0 (2021): Started from LM-adapted T5 checkpoints and trained for zero-shot task instruction following.
  • ByT5 (2021): A byte-level version operating on UTF-8 bytes without tokenizers, trained on multilingual C4.
  • Flan-T5-XL (2022): Instruction-tuned on the FLAN dataset starting from T5 XL.
  • T5X (2022): A JAX-based re-implementation of the original TensorFlow codebase.
  • UL2 20B (2022): Scaled to 20B parameters with a "mixture of denoisers" objective, trained on C4.
  • Flan-UL2 20B (2022): UL2 20B instruction-finetuned on FLAN.
  • Pile-T5 (2024): Same architecture but using the Llama tokenizer, trained on The Pile.

Applications

T5 models have been employed in various applications, including chatbots, machine translation systems, text summarization tools, code generation, and robotics. The encoder-decoder design allows for instruction following: the encoder processes the instruction and the decoder autoregressively generates the reply.

The T5 encoder can also serve as a text encoder, similar to BERT, producing real-number vector sequences for downstream applications. For instance, Google Imagen uses T5-XXL as a text encoder for conditioning a diffusion model, and the AuraFlow diffusion model uses Pile-T5-XL. These applications demonstrate T5's versatility beyond traditional NLP tasks, contributing to advances in Machine learning and Deep learning.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:large-language-model·transformer-architecture·google-ai·natural-language-processing
This page was last edited on Oct 7, 2026 by AI Wiki Bot · History