ByteDance open-sourced Bernini-Diffusers-v2, a video generation and editing model
ByteDance released Bernini-Diffusers-v2 on Hugging Face, an Apache-2.0 licensed image-text-to-video model built on Wan2.2-T2V-A14B and Qwen2.5-VL-7B-Instruct. The release includes inference code and model weights, with a benchmark snapshot showing VBench 84.46 and OpenS2V 63.83.
China context
- Original name
- 字节跳动
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can download the model weights from Hugging Face and use the Apache-2.0 license for commercial applications, with inference code available in the Bernini repository.
- For investors
- ByteDance's release of a full video generation and editing pipeline under Apache-2.0 may intensify competition in the open-weights video model market, potentially affecting valuations of startups in this space.
ByteDance released Bernini-Diffusers-v2 on Hugging Face under Apache-2.0. The model is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer. It packages a Qwen2.5-VL planner, Bernini planning weights, and Wan2.2 diffusion components in a self-contained diffusers-format directory. The release includes inference code and model weights, and supports tasks including t2i, i2i, t2v, v2v, rv2v, and r2v. Benchmark results reported by ByteDance include VBench 84.46, OpenS2V 63.83, EditVerse 8.02, and OpenVE 3.96.
Bernini-Diffusers-v2 uses a training recipe that warms up the connector for thousands of steps before co-training, improving reference-guided video editing and OpenS2V performance compared to the first Bernini-Diffusers release. The model decomposes complex instructions and plans semantic changes before rendering, at the cost of a heavier checkpoint layout than Bernini-R. Recommended environment includes Python 3.11.2, PyTorch 2.7.1+cu126, CUDA 12.6, and Hopper GPUs (H100/H800/H200).
Developers outside China can now use ByteDance's Bernini-Diffusers-v2 under Apache-2.0, gaining access to a video generation and editing pipeline with an MLLM-based semantic planner and Wan2.2 renderer without licensing fees. This may pressure other open-weights video model providers to match its instruction-following and editing capabilities.
Bernini-Diffusers-v2 provides a self-contained diffusers-format directory that can be passed directly to --config, simplifying integration for developers. The Apache-2.0 license permits commercial use, potentially reducing costs for companies needing video generation and editing capabilities.
The next observable signal is whether ByteDance releases a renderer-only Bernini-R version with lighter checkpoint layout, or whether independent benchmarks confirm the reported VBench and OpenS2V scores. Adoption can be tracked via Hugging Face downloads and community fine-tunes.