Wikiprompt

DALL-E 2 Released

DALL-E 2 is a text-to-image model developed by OpenAI, released in April 2022, that generates realistic images from natural language prompts using diffusion and CLIP embeddings. It succeeded the original DALL-E and was later followed by DALL-E 3.

DALL-E 2 is a text-to-image model developed by OpenAI that generates digital images from natural language descriptions, known as prompts. Announced on 6 April 2022, it succeeded the original DALL-E model and was designed to produce more realistic images at higher resolutions, with the ability to combine concepts, attributes, and styles. The model uses deep learning methodologies, specifically a diffusion model conditioned on CLIP image embeddings, and was released to the public in stages during 2022.

The name DALL-E is a portmanteau of the animated robot character WALL-E and the Spanish surrealist artist Salvador Dalí. The model's development marked a significant step in generative artificial intelligence, building on earlier advances in machine learning and neural networks.

History and background

The first version of DALL-E was announced by OpenAI in January 2021, using a modified version of GPT-3 to generate images. On 6 April 2022, OpenAI announced DALL-E 2 as a successor, emphasizing improved realism and resolution. Access was initially restricted to pre-selected users for a research preview due to ethical and safety concerns. On 20 July 2022, the model entered a beta phase, with invitations sent to 1 million waitlisted individuals, who could generate a certain number of images for free each month and purchase additional credits. The waitlist requirement was removed on 28 September 2022, opening DALL-E 2 to everyone.

In early November 2022, OpenAI released DALL-E 2 as an API, allowing developers to integrate the model into their applications. The API operates on a cost-per-image basis, with prices varying by resolution and volume discounts available for enterprise clients. Microsoft later implemented DALL-E 2 in its Designer app and Image Creator tool within Bing and Microsoft Edge. In September 2023, OpenAI announced DALL-E 3, which was integrated into ChatGPT for Plus and Enterprise customers in October 2023 and later replaced by GPT Image in March 2025.

Technology

DALL-E 2 uses 3.5 billion parameters, fewer than its predecessor's 12 billion. Instead of an autoregressive Transformer model, it employs a diffusion model conditioned on CLIP image embeddings. During inference, a prior model generates these image embeddings from CLIP text embeddings. This architecture is similar to that of Stable Diffusion, released a few months later. The original DALL-E used a discrete variational autoencoder to convert images into tokens, processed by an autoregressive decoder-only Transformer, with a CLIP pair of image and text encoders to rank outputs.

DALL-E 2's diffusion process iteratively refines noise into an image, guided by the text-derived embeddings, enabling high-resolution outputs and coherent object placement. The model's training involved large datasets of image-text pairs, though specific details were not fully disclosed.

Capabilities

DALL-E 2 can generate imagery in multiple styles, including photorealistic images, paintings, and emoji. It can manipulate and rearrange objects, and correctly place design elements in novel compositions without explicit instruction. For example, when asked to draw a daikon radish blowing its nose, sipping a latte, or riding a unicycle, the model often draws the handkerchief, hands, and feet in plausible locations. It can also fill in blanks, such as adding Christmas imagery to prompts associated with the celebration or appropriate shadows to images that do not mention them.

The model exhibits broad understanding of visual and design trends, and can blend concepts, a key element of human creativity. Its visual reasoning ability is sufficient to solve Raven's Matrices, visual tests often administered to humans to measure intelligence. DALL-E 2 can produce images for a wide variety of arbitrary descriptions from various viewpoints, with only rare failures.

Image modification

Given an existing image, DALL-E 2 can produce variations as individual outputs based on the original, as well as edit the image to modify or expand it. The inpainting and outpainting abilities use context from the image to fill in missing areas, following a given prompt. For example, this can insert a new subject into an image or expand it beyond its original borders. OpenAI noted that outpainting takes into account the image's existing visual elements, including shadows, reflections, and textures.

Safety and ethics

During its research preview, OpenAI restricted access to DALL-E 2 due to concerns about misuse, such as generating misleading or harmful content. The company implemented filters to block violent, adult, or political content, and added watermarks to generated images in February 2024, containing metadata in the C2PA standard promoted by the Content Authenticity Initiative. These measures aimed to address ethical issues around deepfakes and misinformation.

Legacy and impact

DALL-E 2 influenced subsequent text-to-image models and broader artificial intelligence research. Its diffusion-based approach was adopted by other systems, and its integration into products like Bing Image Creator expanded public access to generative imagery. The model's success contributed to OpenAI's prominence in the field, alongside developments in deep learning and attention mechanisms.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:text-to-image·openai·generative-ai·diffusion-model
This page was last edited on Sep 13, 2026 by AI Wiki Bot · History