让视频去物更智能:同时清除物体及其影响,效果更真实。
EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

- 用视觉语言模型分析物体引发的多种影响,指导精准清除
- 在复杂场景中实现90%以上覆盖,时间连续性保持良好
- 适合处理现实视频中多变、隐蔽的影响,如动态反射和残影
视频物体移除不仅要消除目标物体,还需同步清除其引发的各类影响,同时保持高保真度与时空一致性。现有方法主要依赖预定义效应类别和固定数据分布隐式学习物体-效应对应关系,难以泛化到包含组合效应、空间分离或弱关联效应、长尾物理现象及动态交互的复杂真实场景。我们提出EffectLearner,一个融合视觉语言模型(VLM)物体-效应推理器与扩散图像变换器(DiT)视频擦除器的语义增强框架。在结构化提示引导下,推理器对目标突出的视频进行跨模态推理,提取紧凑的效应感知上下文,指导擦除器完成全面的物体-效应清除。运动感知掩码引导与运动一致性监督进一步提升物体运动及场景动态变化下的覆盖范围与时空稳定性。为充分挖掘该框架在挑战性真实场景中的潜力,我们构建了专门针对复杂物体诱导效应的成对视频数据集EffectWorld,并引入渐进式训练课程,结合常规监督与复杂效应数据。在标准ROSE-Bench上,EffectLearner在多数指标上优于现有基线,在EffectWorld-Eval和极具挑战性的EffectWorld-Wild上均表现显著领先,证明其在复杂真实场景中实现高质量视频物体移除的能力。
原文摘要 · Abstract (English)
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。