无需成对数据,用关键帧控制实现高质量视频编辑
NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
- 用关键帧提供语义引导,动态融合原视频运动纹理
- 在多个基准上超越现有方法,保持动作连贯与画面清晰
- 适合需要快速编辑且无配对数据的创作者
近期视频编辑模型取得显著进展,但大多依赖大规模成对数据。大规模自然对齐数据的收集仍具挑战性,尤其在局部视频编辑中尤为突出。现有方法通过全局运动控制将图像编辑迁移至视频,但难以保证背景一致性和时间连贯性。本文提出 NOVA:稀疏控制与密集合成框架,用于无配对视频编辑。稀疏分支通过用户编辑的关键帧分布提供语义引导,密集分支持续融合原始视频的运动与纹理信息以维持高保真度与连贯性。此外,我们引入退化模拟训练策略,在人工退化视频上训练模型学习运动重建与时间一致性,从而无需配对数据。大量实验表明,NOVA 在编辑保真度、动作保留和时间连贯性方面均优于现有方法。
原文摘要 · Abstract (English)
Recent video editing models have achieved impressive results, but most still require large-scale paired datasets. Collecting such naturally aligned pairs at scale remains highly challenging and constitutes a critical bottleneck, especially for local video editing data. Existing workarounds transfer image editing to video through global motion control for pair-free video editing, but such designs struggle with background and temporal consistency. In this paper, we propose NOVA: Sparse Control \& Dense Synthesis, a new framework for unpaired video editing. Specifically, the sparse branch provides semantic guidance through user-edited keyframes distributed across the video, and the dense branch continuously incorporates motion and texture information from the original video to maintain high fidelity and coherence. Moreover, we introduce a degradation-simulation training strategy that enables the model to learn motion reconstruction and temporal consistency by training on artificially degraded videos, thus eliminating the need for paired data. Our extensive experiments demonstrate that NOVA outperforms existing approaches in edit fidelity, motion preservation, and temporal coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。