用锚点帧+扩散模型,让长视频编辑更连贯稳定。
AnchorSync: Global Consistency Optimization for Long Video Editing
- 分步处理:先编辑稀疏锚点帧,再插值中间帧
- 在分钟级视频上显著减少结构漂移和时间伪影
- 适合需要长期一致性的视频编辑场景
长视频编辑因需保持数千帧间的全局一致性与时间连贯性而极具挑战。现有方法常出现结构漂移或时间伪影,尤其在分钟级序列中更为明显。我们提出AnchorSync,一种基于扩散模型的新框架,通过将任务解耦为稀疏锚点帧编辑与平滑中间帧插值,实现高质量长时视频编辑。该方法通过渐进去噪过程强化结构一致性,并利用多模态引导保留时间动态。大量实验表明,AnchorSync在视觉质量和时间稳定性上均优于先前方法。
原文摘要 · Abstract (English)
Editing long videos remains a challenging task due to the need for maintaining both global consistency and temporal coherence across thousands of frames. Existing methods often suffer from structural drift or temporal artifacts, particularly in minute-long sequences. We introduce AnchorSync, a novel diffusion-based framework that enables high-quality, long-term video editing by decoupling the task into sparse anchor frame editing and smooth intermediate frame interpolation. Our approach enforces structural consistency through a progressive denoising process and preserves temporal dynamics via multimodal guidance. Extensive experiments show that AnchorSync produces coherent, high-fidelity edits, surpassing prior methods in visual quality and temporal stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。