Event date · · Wan Team

Wan: Open and Advanced Large-Scale Video Generative Models

FACT STATEMENT

In March 2025, the Wan team released the open-source video foundation model series Wan, including 1.3B and 14B versions. Based on the diffusion Transformer architecture, Wan consistently outperforms existing open-source models and commercial solutions in internal and external benchmarks. The 14B model leads in multiple downstream tasks (e.g., image-to-video, instruction-guided video editing, personalized video generation). The 1.3B model requires only 8.19GB of VRAM and can run on consumer-grade GPUs. All code and models are open-sourced.

What happened

Wan is the first fully open-source large-scale video generation model series. Its 14B version surpasses existing open-source and commercial models in performance, while the 1.3B version achieves usability on consumer-grade GPUs. Through innovative VAE, scalable pre-training strategies, and large-scale data curation, this work advances the democratization of video generation. The open-source strategy is expected to accelerate community innovation and application deployment.

Technical significance

Wan is based on the diffusion Transformer (DiT) architecture, with core innovations including: 1) a novel VAE that improves video compression and reconstruction quality; 2) a scalable pre-training strategy trained on billions of images and videos, validating scaling laws for video generation; 3) a large-scale data curation method ensuring data quality and diversity; 4) automated evaluation metrics. The model supports 8 downstream tasks, including text-to-video, image-to-video, video editing, and personalized generation. The 14B model achieves SOTA on multiple benchmarks (e.g., UCF-101, MSR-VTT), but the paper does not provide detailed comparison tables. The 1.3B model requires only 8.19GB of VRAM, suitable for consumer GPUs like RTX 4090.

Industry impact

Wan's open-source strategy will significantly lower the barrier to video generation, impacting industries such as film production, advertising, and social media content creation. Existing commercial models (e.g., Runway, Pika) face competitive pressure. Meanwhile, the low resource requirements of the 1.3B model enable small studios and individual creators to use high-quality video generation tools.

Decision value

It is recommended that content creation platforms (e.g., Douyin, YouTube) evaluate integrating Wan models to enhance user creation tools. Investment opportunities exist in vertical applications based on Wan (e.g., advertising video generation, educational video production). Engineering-wise, the 1.3B model can be deployed on edge devices for real-time video generation.

What to watch

Areas to watch: 1) Wan's performance on long video generation (>1 minute); 2) model controllability and consistency; 3) community secondary development and optimization based on the open-source model; 4) copyright and ethical issues (e.g., deepfakes); 5) ongoing comparison with closed-source models like Sora.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.