用统一噪声控制实现精准视频物体编辑,效果超越现有方法。
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
- 引入随机自适应噪声作为统一编辑信号,支持多种操作。
- 在多个任务中实现高保真编辑,人评与自动指标均领先。
- 适合需要高效多类型视频编辑的研究者与开发者。
扩散模型近期推动了视频编辑的发展,但可控编辑仍具挑战,因需精确操控多样物体属性。现有方法对不同编辑任务需不同控制信号,导致模型设计复杂且训练资源消耗大。为此,我们提出O-DisCo-Edit,一个融合新型物体畸变控制(O-DisCo)的统一框架。该信号基于随机与自适应噪声,将多种编辑提示统一编码于单一表示中。结合“复制-重构”保留模块以保护未编辑区域,O-DisCo-Edit通过高效训练范式实现高保真、高效的编辑。大量实验与全面的人工评估一致表明,O-DisCo-Edit在多种视频编辑任务中均优于专用及多任务最先进方法。
原文摘要 · Abstract (English)
Diffusion models have recently advanced video editing, yet controllable editing remains challenging due to the need for precise manipulation of diverse object properties. Current methods require different control signal for diverse editing tasks, which complicates model design and demands significant training resources. To address this, we propose O-DisCo-Edit, a unified framework that incorporates a novel object distortion control (O-DisCo). This signal, based on random and adaptive noise, flexibly encapsulates a wide range of editing cues within a single representation. Paired with a "copy-form" preservation module for preserving non-edited regions, O-DisCo-Edit enables efficient, high-fidelity editing through an effective training paradigm. Extensive experiments and comprehensive human evaluations consistently demonstrate that O-DisCo-Edit surpasses both specialized and multitask state-of-the-art methods across various video editing tasks. https://cyqii.github.io/O-DisCo-Edit.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。