让视频删物更真实,能自动修复物体互动的物理错误。
VOID: Video Object and Interaction Deletion
- 用因果推理识别被删物体影响的区域,引导生成合理替代画面。
- 在合成与真实数据上,动态一致性优于现有方法。
- 适合需要真实物理模拟的视频编辑与影视制作人员。
现有视频物体移除方法擅长修复物体背后的背景内容及外观伪影(如阴影、反光),但在处理物体间复杂交互(如碰撞)时,无法修正其下游物理影响,导致结果不真实。本文提出VOID框架,专为复杂场景下的物理合理补全设计。通过Kubric与HUMOTO生成新的成对反事实移除数据集,其中移除物体需改变后续物理互动。推理时,视觉-语言模型识别受影响区域,并引导视频扩散模型生成符合物理规律的替代结果。在合成与真实数据上的实验表明,该方法在物体移除后更有效地保持了场景动态的一致性,优于以往方法。我们希望此框架推动视频编辑模型向世界级模拟器演进,实现高层次因果推理。
原文摘要 · Abstract (English)
Existing video object removal methods excel at inpainting content "behind" the object and correcting appearance-level artifacts such as shadows and reflections. However, when the removed object has more significant interactions, such as collisions with other objects, current models fail to correct them and produce implausible results. We present VOID, a video object removal framework designed to perform physically-plausible inpainting in these complex scenarios. To train the model, we generate a new paired dataset of counterfactual object removals using Kubric and HUMOTO, where removing an object requires altering downstream physical interactions. During inference, a vision-language model identifies regions of the scene affected by the removed object. These regions are then used to guide a video diffusion model that generates physically consistent counterfactual outcomes. Experiments on both synthetic and real data show that our approach better preserves consistent scene dynamics after object removal compared to prior video object removal methods. We hope this framework sheds light on how to make video editing models better simulators of the world through high-level causal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。