Tencent Hunyuan open-sources Prism, a dynamic sparse attention framework for native 2K joint video-audio generation
Tencent Hunyuan released Prism, an open-source dynamic sparse attention framework for natively training joint video-audio generation models at 2K resolution. The repository includes training/inference code, a technical report, and a preview model checkpoint supporting 720p, 1080p, and 2K. Prism achieves a 2.5× training speedup over full attention while surpassing it in generation quality.
China context
- Original name
- 腾讯混元
- Outside China
- Open weights · github.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers can access the training and inference code, along with a preview model checkpoint, to experiment with native joint video-audio generation at up to 2K resolution without incurring full attention's quadratic cost.
- For investors
- The release signals Tencent Hunyuan's commitment to open-source multimodal research, influencing the competitive landscape for video generation startups by lowering the barrier to high-resolution model development.
Tencent Hunyuan, in collaboration with Fudan University and Zhejiang University, open-sourced Prism, a dynamic sparse attention framework designed for natively training joint video-audio generation models at 2K resolution. Prism organizes token sequences into spatiotemporal macro-zones and dynamically assigns tailored block shapes based on visual feature variance and audio-to-video cross-attention norms, enabling efficient training while preserving cross-modal interactions. The repository provides training and inference code, a technical report, and a preview model checkpoint that supports native joint video-audio generation at 720p, 1080p, and 2K resolutions. According to the project, Prism achieves a 2.5× training speedup compared to full attention and surpasses it in generation quality.
Prism introduces a dynamic sparse attention mechanism that partitions token sequences into spatiotemporal macro-zones. For each zone, it estimates local information structure using video feature variance along channels and feature norms from audio-to-video cross-attention, then assigns a tailored block shape with finer partitioning along axes of rapid visual variation and strong audio-visual coupling. A hybrid block selection strategy dynamically determines per-query sparsity. This approach reduces quadratic attention cost while maintaining semantically coherent blocks that capture both visual content and joint video-audio interaction patterns.
Developers outside China can now access a training-efficient framework for high-resolution joint video-audio generation, reducing compute costs for building multimodal models. The open-source release of training code and a preview checkpoint lowers the barrier for researchers and startups to experiment with 2K video-audio generation without full attention's quadratic overhead.
For organizations building video generation products, Prism offers a path to train models at higher resolutions with reduced computational cost, enabling more detailed and dynamic video-audio content. The open-source availability allows integration into existing pipelines without licensing fees.
The project's to-do list indicates a future release of 'Prism-pro', suggesting ongoing development beyond the current preview checkpoints. Verification of the claimed 2.5× training speedup and generation quality improvements would require independent benchmarking against full attention baselines.