DALL-E 3 is a text-to-image model developed by OpenAI, released in October 2023. It is the third iteration in the DALL-E series, following DALL-E and DALL-E 2. DALL-E 3 uses deep learning methodologies to generate digital images from natural language descriptions, known as prompts, with a focus on understanding "significantly more nuance and detail" than its predecessors. The model was integrated natively into ChatGPT for ChatGPT Plus and ChatGPT Enterprise customers, and later made available via OpenAI's API and Labs platform. In March 2025, DALL-E 3 was replaced in ChatGPT by GPT Image's native image-generation capabilities.
History and background
DALL-E was first announced by OpenAI on 5 January 2021, using a modified version of GPT-3 to generate images. DALL-E 2 followed on 6 April 2022, offering higher resolution and the ability to combine concepts, attributes, and styles. After a beta phase in July 2022 and a public release in September 2022, DALL-E 2 was made available via API in November 2022. In September 2023, OpenAI announced DALL-E 3, which was released in October 2023. Microsoft implemented DALL-E 3 in Bing's Image Creator tool, and Microsoft Copilot runs on DALL-E 3. The name "DALL-E" is a portmanteau of the Pixar robot WALL-E and the Spanish surrealist artist Salvador DalÃ. In February 2024, OpenAI began adding watermarks to DALL-E generated images, using metadata in the C2PA standard.
Technology
DALL-E 3 builds on the transformer architecture used in OpenAI's GPT models. While a technical report was written for DALL-E 3, it does not include training or implementation details, instead focusing on improved prompt following. The model is designed to interpret complex prompts with greater accuracy and detail, and can generate more coherent and accurate text within images compared to earlier versions.
DALL-E and DALL-E 2
The original DALL-E used a discrete VAE, an autoregressive decoder-only transformer with 12 billion parameters, and a CLIP model for ranking outputs. DALL-E 2, with 3.5 billion parameters, shifted to a diffusion model conditioned on CLIP image embeddings, similar to the architecture later used in Stable Diffusion.
Capabilities
DALL-E 3 can generate imagery in multiple styles, including photorealistic images, paintings, and emoji. It can manipulate and rearrange objects, and place design elements in novel compositions. It follows complex prompts with more accuracy and detail than its predecessors, and can produce coherent text within images. DALL-E 3 is integrated into ChatGPT Plus, allowing users to generate images directly in conversations.
Image modification
DALL-E 2 and DALL-E 3 can produce variations of existing images and edit them, including inpainting and outpainting. These features use context from the original image to fill in missing areas or expand beyond borders, maintaining consistency with shadows, reflections, and textures.
Reception and impact
DALL-E 3 has been noted for its enhanced prompt following and integration with ChatGPT, making it accessible to a broader audience. Its use in Microsoft's Bing Image Creator and Copilot has expanded its reach. The model's ability to generate detailed and accurate images from complex prompts has influenced the field of generative AI, though concerns about ethics and safety have been raised, leading to measures like watermarks.