One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
EditVid is a training-free video editing framework that combines sparse causal memory, correspondence-based post-attention token injection, and soft latent blending. It supports instruction-guided and reference-guided edits including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc compared to 58.95 for the strongest evaluated training-free baseline, and obtains competitive results on IVEBench. A user study shows 51.8% overall preference for EditVid over 7 competing methods.
EditVid is a training-free framework for diverse video editing that unifies instruction-guided and reference-guided editing. It uses sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The framework supports style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, outperforming the strongest training-free baseline (58.95), and shows competitive results on IVEBench. A user study indicates 51.8% overall preference for EditVid over 7 competing methods.
EditVid's training-free design leverages sparse causal memory to maintain local temporal coherence, correspondence-based post-attention token injection to preserve identity across frames, and soft latent blending to localize edits. This combination enables a single framework to handle both instruction-guided and reference-guided editing tasks without task-specific fine-tuning, achieving significant improvements on FiVE-Acc.
The unified training-free approach reduces the need for task-specific models and large-scale training, potentially lowering development costs and enabling faster deployment for video editing applications. The strong performance on FiVE and competitive results on IVEBench suggest EditVid could become a practical baseline for diverse video editing tasks in creative tools and media production.
EditVid's training-free nature and unified framework could reduce computational overhead and simplify integration into existing video editing workflows, offering potential cost savings and faster iteration for companies developing AI-assisted video editing tools.
Future work may focus on extending EditVid to longer videos, improving real-time performance, and validating on additional benchmarks. Adoption by video editing software vendors or integration into generative AI pipelines could be observed through product announcements or further research citations.