GPT Image is a family of Text-to-image generation generation models developed by OpenAI. It replaced the earlier DALL-E series as OpenAI's image generation lineage, and unlike DALL-E it is natively multimodal: image generation is performed by the same model family that powers ChatGPT, rather than by a separate diffusion model called through a bridge prompt.
Members
- GPT Image 1 (April 2025). The first API release of the technology behind ChatGPT's native image generation, which had gone viral weeks earlier through the Studio Ghibli-style portrait wave. Established the family's signature strengths: legible in-image text and strong instruction following.
- GPT Image 1.5 (late 2025). An interim upgrade with better editing precision and faster generation.
- GPT Image 2 (2026). The current flagship, with improved photorealism, longer-text rendering and identity-preserving editing.
Significance
The family's native multimodality changed prompt-writing practice. Because the generator understands conversational context and structured briefs directly, prompt engineers moved from keyword lists (the diffusion-era style still typical of Stable Diffusion checkpoints) toward long-form natural language briefs with explicit constraints, a style that now dominates prompt-sharing platforms. On Wikiprompt, GPT Image family models together account for the largest share of image prompts in the catalog.
See also
- DALL-E, the predecessor series
- Nano Banana, Google's competing image model line
- Midjourney