无需训练即可精准编辑视频与3D场景,保留原内容同时实现复杂修改。
V2Edit: Versatile Video Diffusion Editor for Videos and 3D Scenes
- 分步分解编辑任务,通过噪声与文本注意力协同控制
- 在多个视频与3D场景任务中达到当前最优效果
- 适合需要高质量、几何一致编辑的影视与动画创作者
本文提出V²Edit,一种无需训练的指令引导视频与3D场景编辑框架。针对保持原始内容与完成编辑任务之间的平衡难题,该方法采用渐进式策略,将复杂编辑任务拆分为一系列简单子任务。每个子任务通过初始噪声、去噪过程中的噪声添加以及文本提示与视频内容间的交叉注意力图三个协同机制进行控制,确保原始视频元素的有效保留,同时准确执行目标编辑。在原有视频编辑能力基础上,通过“渲染-编辑-重建”流程拓展至3D场景编辑,支持包括物体插入在内的显著几何变化,实现高质量且三维一致的编辑结果。大量实验表明,V²Edit在多种挑战性视频编辑和复杂3D场景编辑任务中均取得优异表现,成为两个领域的最新标杆。
原文摘要 · Abstract (English)
This paper introduces V$^2$Edit, a novel training-free framework for instruction-guided video and 3D scene editing. Addressing the critical challenge of balancing original content preservation with editing task fulfillment, our approach employs a progressive strategy that decomposes complex editing tasks into a sequence of simpler subtasks. Each subtask is controlled through three key synergistic mechanisms: the initial noise, noise added at each denoising step, and cross-attention maps between text prompts and video content. This ensures robust preservation of original video elements while effectively applying the desired edits. Beyond its native video editing capability, we extend V$^2$Edit to 3D scene editing via a "render-edit-reconstruct" process, enabling high-quality, 3D-consistent edits even for tasks involving substantial geometric changes such as object insertion. Extensive experiments demonstrate that our V$^2$Edit achieves high-quality and successful edits across various challenging video editing tasks and complex 3D scene editing tasks, thereby establishing state-of-the-art performance in both domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。