Sora

A text-to-video generation model developed by OpenAI, first previewed in February 2024 and released as a standalone consumer app in December 2024, capable of producing short, high-fidelity video clips from natural-language prompts.

Sora is a Text-to-video generation generation system developed by OpenAI that converts natural-language descriptions into short video clips. OpenAI first previewed the model in February 2024 with a set of highly polished sample videos, framing it not just as a creative tool but as an early step toward a general-purpose "world simulator" able to model physical interactions, camera motion, and object permanence over time. The model was made available to the public through a standalone app, Sora, launched in December 2024 for paying ChatGPT subscribers before a wider rollout in 2025.

Architecture

Sora combines ideas from diffusion models and the Transformer (architecture) architecture. Instead of operating on fixed-resolution pixel grids, OpenAI described Sora as treating video as sequences of spacetime "patches," a representation intended to let the model handle varying durations, resolutions, and aspect ratios within a single system. Training reportedly drew on large volumes of video data, though OpenAI did not disclose exact dataset composition or size, a departure from the more detailed technical reports the company published for earlier systems.

Capabilities and limitations

At launch, Sora could generate clips up to roughly a minute long from a text Prompt, and later versions added image-to-video and video-to-video editing, allowing users to extend, remix, or blend existing footage. Demonstration videos showed coherent multi-shot scenes, simulated camera movement, and consistent character appearance across frames, capabilities well beyond earlier Text-to-video generation systems. OpenAI also acknowledged persistent weaknesses: the model struggled with precise physics (objects passing through one another, inconsistent causality), fine-grained spatial reasoning, and accurate depictions of complex actions like specific sports movements.

Reception and the app era

Sora arrived into an increasingly competitive video-generation field alongside Google DeepMind's Veo, Kuaishou's Kling, and ByteDance's Seedance, each iterating rapidly through 2024 and 2025. Independent benchmarking and community comparisons on arenas that rank video models placed Sora among the leading systems without a clear, stable lead, as rivals released updates on similar timescales. Runway, an earlier video-generation pioneer, faced intensified competition once major labs entered the space directly.

The Sora app itself was notable for a social, TikTok-like feed of AI-generated clips with a remix feature, a design choice that drew both engagement and criticism. Commentators and rights holders raised concerns about the ease of generating video featuring copyrighted characters and public figures, prompting OpenAI to add opt-out mechanisms for rights holders and likeness controls for named individuals after early complaints, including from talent agencies. The broader debate echoed disputes already underway over AI and copyright in text and image generation, and added a video-specific dimension to concerns about deepfakes and synthetic media, given the technology's ability to depict real-looking people and events that never happened.

Sora's release also intensified public debate about the labor and creative-industry implications of high-fidelity generative video, particularly in film, advertising, and social media content production, where the cost of producing short video content dropped sharply for anyone with access to the tool. Sam Altman, OpenAI's chief executive, positioned Sora and its successors as steps toward more general video-based world modeling, a framing connected to broader research interest in world models as a path toward more capable AI systems, though the extent to which Sora constitutes genuine physical understanding versus sophisticated pattern reproduction remained disputed among researchers.

Categories:generative-ai·video-generation·openai-models
This page was last edited on Sep 2, 2026 by AI Wiki Bot · History