让视频编辑同时对齐音画,精确到具体物体
AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
- 用可变粒度掩码精修器,把粗略选择转为精准目标区域
- 自反馈音频代理生成高质量时序指引,实现音画同步控制
- 专为物体级编辑设计,适合需要精细音画协同的创作者
近期视频生成进展表明,逼真的音画同步对吸引人的内容创作至关重要。然而,现有视频编辑方法大多忽视音画同步,且缺乏精确的时空可控性以实现实例级编辑。本文提出 AVI-Edit,一种用于音画同步的视频实例编辑框架。我们设计了一个粒度感知的掩码精修器,可将用户提供的粗略掩码迭代优化为精确的实例级区域。同时,引入自反馈音频代理,生成高质量音频引导信号,提供细粒度的时间控制能力。为支持该任务,我们还构建了一个大规模数据集,包含以实例为中心的对应关系和全面标注。大量实验证明,AVI-Edit 在视觉质量、条件遵循度和音画同步方面均优于当前最佳方法。
原文摘要 · Abstract (English)
Recent advancements in video generation highlight that realistic audio-visual synchronization is crucial for engaging content creation. However, existing video editing methods largely overlook audio-visual synchronization and lack the fine-grained spatial and temporal controllability required for precise instance-level edits. In this paper, we propose AVI-Edit, a framework for audio-sync video instance editing. We propose a granularity-aware mask refiner that iteratively refines coarse user-provided masks into precise instance-level regions. We further design a self-feedback audio agent to curate high-quality audio guidance, providing fine-grained temporal control. To facilitate this task, we additionally construct a large-scale dataset with instance-centric correspondence and comprehensive annotations. Extensive experiments demonstrate that AVI-Edit outperforms state-of-the-art methods in visual quality, condition following, and audio-visual synchronization. Project page: https://hjzheng.net/projects/AVI-Edit/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。