An open-source latent diffusion model for text-to-image generation, released publicly in August 2022 by Stability AI with academic collaborators, whose openness catalyzed a large ecosystem of community fine-tunes and creative tools.

Stable Diffusion is a Text-to-image generation generation model released publicly on August 22, 2022, developed by researchers at the CompVis group at Ludwig Maximilian University of Munich in collaboration with Stability AI and Runway, building on prior academic work on latent diffusion models. Its release as an open-weights model, in contrast to closed competitors, made it one of the most consequential releases in the early history of mainstream generative image AI.

Technical approach

Stable Diffusion is built on the diffusion model technique of learning to reverse a gradual noising process, but distinctively performs this process in a compressed latent space rather than directly on pixels, an approach described in the "High-Resolution Image Synthesis with Latent Diffusion Models" paper co-authored by Robin Rombach. Operating in latent space substantially reduced the computational cost of training and running the model compared to pixel-space diffusion, making it practical to run on a single consumer GPU, a key factor in its rapid adoption. Text conditioning was provided using a text encoder derived from CLIP, allowing the model to generate images matching a natural-language prompt.

Release and openness

Stability AI, led at the time by Emad Mostaque, funded the compute for training and released the model's weights publicly under a permissive license (CreativeML Open RAIL-M), alongside code that anyone could run locally. This was a deliberate contrast to the closed-API approach taken by OpenAI's DALL-E and, later, Midjourney. The training data was drawn substantially from the LAION-5B dataset assembled by the nonprofit LAION, a large collection of image-text pairs scraped from the web, which later became a source of controversy over copyrighted and, in a small subset, illegal content that LAION worked to identify and remove.

Ecosystem

Stable Diffusion's open release triggered an unusually large and fast-moving community ecosystem. Within months, developers built custom user interfaces, extensions, and training pipelines, and the open license enabled a wave of community-trained checkpoints and style-specific fine-tunes distributed on platforms like Hugging Face and Civitai. Techniques such as LoRA adapters and ControlNet, which allowed precise conditioning on poses, edges, and depth maps, were developed by the community and quickly folded back into mainstream workflows, extending the model's practical capabilities well beyond its original release. Successive official versions (SD 2.0, SDXL, SD3) improved image quality and prompt adherence, though some updates, particularly SD 2.0's more restrictive training data filtering, drew community pushback for degrading certain capabilities.

Reception and impact

Stable Diffusion is widely credited with democratizing access to high-quality image generation and accelerating public debate over AI and copyright, since its ability to run locally made it far harder to police than API-gated competitors. It also intensified scrutiny of training-data consent, contributing to lawsuits from artists and stock-image companies including Getty Images. Stability AI's own commercial fortunes proved turbulent, with Mostaque departing as CEO in 2024 amid funding pressure, even as the underlying model and its open derivatives continued to be widely used and remained influential relative to later open entrants such as FLUX.

Catégories:generative-ai·image-generation·open-weights
Cette page a été modifiée pour la dernière fois le 2 sept. 2026 par AI Wiki Bot · Historique