arXiv:2608.20107cs.CV2026-08

提出新基准与评估方法,让视频去物更符合真实物理因果。

BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal

论文配图:BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal
图 1 · 摘自论文原文
  • 构建含真实物理效应的视频去物配对数据集,支持掩码和指令编辑。
  • 发现现有模型虽保真度高,却无法消除阴影、反射等次生效应。
  • 新评估协议更贴近人类判断,推动视频生成向因果一致性演进。

生成式视频模型在视频物体移除方面显著提升了视觉真实感,但现有评估仍聚焦掩码区域保真度,将移除视为局部修复。现实中,物体移除是因果干预:需同时消除其引发的物理效应,如阴影、反光、光照变化、透光性及动态痕迹。现有基准缺乏对齐的干净参考或局限于简化合成场景,难以系统评估因果一致性。本文提出BeyondMasks,一个包含时序对齐的合成与真实视频对的成对基准,配有干净背景参考。数据集涵盖多样的光照、几何、体积与动态交互,支持基于掩码和指令的编辑。进一步提出CORE——一种基于结构化视觉语言模型的评估协议,联合衡量物体消失与后效一致性,比现有指标更贴近人类判断。对前沿方法的基准测试显示,尽管掩码区域保真度高,但对次生物理效应的消除存在系统性失败,暴露出视觉合理性与因果正确性之间的差距。BeyondMasks将视频物体移除重定义为因果场景一致性问题,并提供统一评估框架。

原文摘要 · Abstract (English)

Recent advances in generative video models have significantly improved visual realism in video object removal, yet evaluation protocols still focus on masked region fidelity, treating removal as local inpainting. In real scenes, object removal is a causal intervention: eliminating an object also requires removing its induced physical effects, such as shadows, reflections, illumination changes, translucency, and dynamic traces. Existing benchmarks lack aligned clean references or remain limited to simplified synthetic settings, preventing systematic evaluation of causal consistency. We introduce BeyondMasks, a paired benchmark for causally consistent video object removal, consisting of temporally aligned synthetic and real world video pairs with clean background references. The dataset spans diverse photometric, geometric, volumetric, and dynamic interactions, and supports both mask based and instruction driven editing. We further propose CORE, a structured vision language model based evaluation protocol that jointly measures object disappearance and after effect consistency, aligning more closely with human judgments than existing metrics. Benchmarking state of the art methods reveals systematic failures in removing secondary physical effects despite high masked region fidelity, exposing a gap between visual plausibility and causal correctness. BeyondMasks reframes video object removal as causal scene consistency rather than local reconstruction and provides a unified framework for its evaluation.

视频生成因果一致性评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。