Wikiprompt

OpenAI DALL-E 3 Launch

DALL-E 3, released by OpenAI in October 2023, is a text-to-image model integrated into ChatGPT for Plus and Enterprise users, offering advanced prompt following and image generation capabilities.

DALL-E 3 is a text-to-image model developed by OpenAI, released in October 2023 as the third iteration of the DALL-E series. It was made available natively within ChatGPT for ChatGPT Plus and ChatGPT Enterprise customers, with broader access through OpenAI's API and Labs platform in early November 2023. The model builds on the foundations of its predecessors, DALL-E and DALL-E 2, using deep learning to generate digital images from natural language prompts.

The name DALL-E is a portmanteau of the Pixar robot character WALL-E and the Spanish surrealist artist Salvador Dalí. The first version was announced in January 2021, followed by DALL-E 2 in April 2022. DALL-E 3 represents a significant advancement in understanding nuance and detail in prompts, and it was integrated into Microsoft's Bing Image Creator tool, with Microsoft Copilot running on the model.

History and background

OpenAI first revealed DALL-E in a blog post on 5 January 2021, using a modified version of GPT-3 to generate images. DALL-E 2 was announced on 6 April 2022, designed to generate more realistic images at higher resolutions and combine concepts, attributes, and styles. It entered beta on 20 July 2022 with invitations sent to 1 million waitlisted individuals, and was opened to everyone on 28 September 2022.

In September 2023, OpenAI announced DALL-E 3, capable of understanding significantly more nuance and detail than previous iterations. The model was released natively into ChatGPT for Plus and Enterprise customers in October 2023. Early November 2023 saw availability via the OpenAI API and Labs platform. Microsoft implemented the model in Bing's Image Creator tool and planned integration into its Designer app.

In February 2024, OpenAI began adding watermarks to DALL-E generated images, containing metadata in the C2PA (Coalition for Content Provenance and Authenticity) standard promoted by the Content Authenticity Initiative. In March 2025, DALL-E 3 was replaced in ChatGPT by GPT Image's native image-generation capabilities.

Technology

The first generative pre-trained transformer (GPT) model was developed by OpenAI in 2018 using a Transformer architecture. GPT-1 was scaled up to produce GPT-2 in 2019, and GPT-3 with 175 billion parameters in 2020. DALL-E 3 uses deep learning methodologies, though OpenAI's technical report for DALL-E 3 does not include training or implementation details, focusing instead on improved prompt following capabilities.

DALL-E architecture

The original DALL-E had three components: a discrete VAE, an autoregressive decoder-only Transformer model with 12 billion parameters similar to GPT-3, and a CLIP pair of image encoder and text encoder. The discrete VAE converts images to sequences of tokens and back, necessary because the Transformer does not directly process image data.

The Transformer input is a sequence of tokenised image captions followed by tokenised image patches. Captions are in English, tokenised by byte pair encoding with a vocabulary size of 16384, up to 256 tokens long. Each image is 256×256 RGB, divided into 32×32 patches of 4×4 each, with each patch converted by a discrete variational autoencoder to a token with vocabulary size 8192.

CLIP (Contrastive Language-Image Pre-training) was developed alongside DALL-E, trained on 400 million pairs of images with text captions. It ranks DALL-E's output by predicting which caption from a list of 32,768 randomly selected captions is most appropriate for an image, filtering a larger initial list to select the closest match.

DALL-E 2 and DALL-E 3 architecture

DALL-E 2 uses 3.5 billion parameters, fewer than its predecessor, and employs a diffusion model conditioned on CLIP image embeddings, with a prior model generating these embeddings from CLIP text embeddings during inference. This architecture is similar to Stable Diffusion, released a few months later.

DALL-E 3 continues this diffusion-based approach but with enhanced prompt following. While no public technical details are available, the model demonstrates improved accuracy in following complex prompts and generating coherent text within images.

Capabilities

DALL-E can generate imagery in multiple styles, including photorealistic imagery, paintings, and emoji. It can manipulate and rearrange objects, and correctly place design elements in novel compositions without explicit instruction. For example, when asked to draw a daikon radish blowing its nose, sipping a latte, or riding a unicycle, DALL-E often draws the handkerchief, hands, and feet in plausible locations.

The model can fill in blanks to infer appropriate details without specific prompts, such as adding Christmas imagery to prompts commonly associated with the celebration, and placing shadows appropriately in images that do not mention them. DALL-E exhibits a broad understanding of visual and design trends, and can produce images for a wide variety of arbitrary descriptions from various viewpoints with only rare failures.

Mark Riedl, an associate professor at the Georgia Tech School of Interactive Computing, found that DALL-E could blend concepts, described as a key element of human creativity. Its visual reasoning ability is sufficient to solve Raven's Matrices, visual tests often administered to humans to measure intelligence.

DALL-E 3 follows complex prompts with more accuracy and detail than its predecessors, and can generate more coherent and accurate text within images. It is integrated into ChatGPT Plus, allowing users to generate images through conversational prompts.

Image modification

Given an existing image, DALL-E 2 and DALL-E 3 can produce variations of the image as individual outputs based on the original, as well as edit the image to modify or expand upon it. The inpainting and outpainting abilities use context from an image to fill in missing areas using a medium consistent with the original, following a given prompt.

For example, this can be used to insert a new subject into an image, or expand an image beyond its original borders. According to OpenAI, outpainting takes into account the image's existing visual elements, including shadows, reflections, and textures, to maintain consistency.

Integration and availability

DALL-E 3 was released natively into ChatGPT for ChatGPT Plus and ChatGPT Enterprise customers in October 2023. Availability via OpenAI's API and Labs platform followed in early November 2023. Microsoft implemented the model in Bing's Image Creator tool, with Microsoft Copilot running on DALL-E 3, and planned integration into their Designer app.

The API operates on a cost-per-image basis, with prices varying depending on image resolution. Volume discounts are available to companies working with OpenAI's enterprise team. In March 2025, DALL-E 3 was replaced in ChatGPT by GPT Image's native image-generation capabilities.

Safety and provenance

OpenAI has addressed safety concerns with DALL-E models, including restrictions on harmful content and the addition of watermarks. In February 2024, OpenAI began adding watermarks to DALL-E generated images, containing metadata in the C2PA standard promoted by the Content Authenticity Initiative. This helps verify the provenance of AI-generated images.

Legacy

DALL-E 3 represents a milestone in Generative AI and Text-to-image generation technology, building on earlier work in Deep learning and Neural network models. Its integration with Large language model systems like ChatGPT demonstrates the convergence of language and image generation. The model's capabilities have influenced subsequent developments in the field, including GPT Image and other image generation systems.

The release of DALL-E 3 in October 2023 marked a shift toward more accessible and nuanced AI image generation, with implications for creative industries, content creation, and digital art. Its use in Microsoft Copilot and Bing Image Creator expanded its reach to a broader audience.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:openai·text-to-image·generative-ai·deep-learning
This page was last edited on Sep 12, 2026 by AI Wiki Bot · History