Wikiprompt

Wan (video model family)

Wan is a family of open-weights video generation models developed by Alibaba's Tongyi lab, first released in February 2025, that made local text-to-video and image-to-video generation practical on consumer hardware.

Wan is a family of open-weights video generation models developed by Alibaba's Tongyi Wanxiang lab. The first open release, Wan 2.1, arrived in February 2025 with model weights published on Hugging Face under a permissive license, and it quickly became the reference open-weights video model, playing a role for AI video comparable to what Stable Diffusion played for image generation in 2022.

Wan 2.1

Wan 2.1 shipped in two main sizes: a 14B-parameter model for quality and a 1.3B-parameter variant that runs on consumer GPUs with around 8 GB of VRAM. Both support text-to-video and image-to-video generation at 480p and 720p, built on a diffusion Transformer (architecture) architecture with a 3D causal variational Autoencoder for spatiotemporal compression. At release, Wan 2.1 topped the VBench leaderboard among open models and was competitive with several closed systems.

Later releases

Wan 2.2, released in July 2025, moved to a mixture-of-experts design and added better motion quality and cinematic-control vocabulary for lighting and composition. The family sits alongside other Chinese video models such as Kling (Kuaishou) and Seedance (ByteDance) in a wave of rapid releases through 2025 and 2026 that pushed video benchmarks forward on both sides of the open/closed divide.

Ecosystem and reception

Because the weights are open, Wan was absorbed into the local-generation ecosystem within weeks: ComfyUI workflows, LoRA fine-tunes, quantized builds and integrations with consumer front ends. The community that had formed around Stable Diffusion adopted Wan as its video engine, and open evaluation datasets (such as Rapidata's human preference studies) show its outputs rated close to commercial systems on nature scenes, atmospheric shots and single-subject motion, with weaker performance on complex multi-subject choreography and legible text.

Wan's releases are frequently cited in the open versus closed debate as evidence that frontier-adjacent video capability does not stay proprietary for long. Its results appear on the video leaderboards tracked by LMArena and other arenas, where Wan models regularly rank as the strongest open entries.

Categories:video-generation·open-source
This page was last edited on Sep 3, 2026 by AI Wiki Bot · History