同时清除视频物体与复杂效果,提升真实场景处理能力。
DualEraser: Joint Video Object and Effect Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver

- 用双模态提示和解耦专家结构,解决语义与像素的冲突问题。
- 在ROSE和VOR-Eval上分别提升2.16dB和1.44dB,性能领先。
- 适合需要高精度视频编辑的开发者与多媒体应用研究者。
视频物体去除常难以在真实场景中同时消除目标物体及其伴随的复杂物理效应(如烟雾、光效)。我们将其归因于语义-像素的根本性冲突,体现在条件层面的模态不一致和优化层面的目标纠缠。针对条件问题,提出双模态文本提示与多条件能力激发机制,显式注入效应语义并利用多模态先验弥补单一模态信息不足。针对优化问题,设计可学习深度CFG融合模块,自适应平衡文本与掩码条件的主导关系。最后,采用解耦专家架构,由定位器负责语义擦除,保全局器负责像素对齐,打破目标纠缠。大量实验表明,DualEraser在标准基准上达到顶尖性能(如在ROSE和VOR-Eval上分别获得2.16dB和1.44dB增益),并可在开放世界视频中稳健移除复杂效应。
原文摘要 · Abstract (English)
Video object removal frequently struggles to eliminate target objects and their associated complex physical effects (e.g., smoke and light) in real-world scenes. We attribute this challenge to a fundamental semantic--pixel conflict, which manifests at two aspects: condition-level modality dissonance and optimization-level objective entanglement. In terms of conditioning, modality dissonance emerges from single-modality information incompleteness and cross-modal dominance imbalance. During optimization, two conflicting objectives---high-level semantic erasure and pixel-level background preservation---are inextricably entangled within a single model. To address these conflicts, we propose DualEraser, a novel framework for joint video object and effect removal. First, a Bipartite Text prompt and a Multi-Conditional Capability Elicitation (MCCE) mechanism explicitly inject effect semantics and further leverage multimodal priors to address the limitations of individual modalities. Second, a Learnable Deep CFG Fusion (LD-CFG) module adaptively balances the relative dominance between the text and mask conditions. Finally, we introduce a decoupled expert architecture comprising a Locator for semantic erasure and a Preserver for background pixel alignment to break the objective entanglement. Extensive experiments demonstrate that DualEraser achieves state-of-the-art quantitative performance on standard benchmarks (e.g., gains of 2.16 dB and 1.44 dB on ROSE and VOR-Eval, respectively), while enabling robust removal of complex effects in open-world videos. https://cyqii.github.io/DualEraser.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。