arXiv:2604.05898cs.CV2026-04被引 1

构建物理感知视频去物基准,评估移除物体后光影一致性。

Physics-Aware Video Instance Removal Benchmark

  • 设计包含95段视频的物理感知评测集,区分简单与复杂交互场景。
  • 四类方法中PISCO-Removal和UniVideo表现最优,但复杂光照仍难恢复。
  • 首次分离语义、视觉、空间三类错误,揭示真实物理因果缺失问题。

视频实例移除(VIR)需在移除目标物体的同时保持背景完整性和物理一致性,如镜面反射与光照交互。尽管文本引导编辑取得进展,现有基准多关注视觉合理性,常忽略物体移除引发的物理因果效应(如残留阴影)。我们提出物理感知视频实例移除(PVIR)基准,包含95个高质量视频,配有实例级掩码与移除提示。该数据集分为简单与困难子集,后者专门针对复杂物理交互。采用解耦的人类评估协议,从指令遵循、渲染质量、编辑排他性三个维度评估四种代表性方法:PISCO-Removal、UniVideo、DiffuEraser和CoCoCo。结果表明,PISCO-Removal与UniVideo达到当前最佳性能,而DiffuEraser频繁引入模糊伪影,CoCoCo在指令遵循上表现显著不足。在困难子集上持续存在的性能下降,凸显恢复复杂物理副作用的挑战。

原文摘要 · Abstract (English)

Video Instance Removal (VIR) requires removing target objects while maintaining background integrity and physical consistency, such as specular reflections and illumination interactions. Despite advancements in text-guided editing, current benchmarks primarily assess visual plausibility, often overlooking the physical causalities, such as lingering shadows, triggered by object removal. We introduce the Physics-Aware Video Instance Removal (PVIR) benchmark, featuring 95 high-quality videos annotated with instance-accurate masks and removal prompts. PVIR is partitioned into Simple and Hard subsets, the latter explicitly targeting complex physical interactions. We evaluate four representative methods, PISCO-Removal, UniVideo, DiffuEraser, and CoCoCo, using a decoupled human evaluation protocol across three dimensions to isolate semantic, visual, and spatial failures: instruction following, rendering quality, and edit exclusivity. Our results show that PISCO-Removal and UniVideo achieve state-of-the-art performance, while DiffuEraser frequently introduces blurring artifacts and CoCoCo struggles significantly with instruction following. The persistent performance drop on the Hard subset highlights the ongoing challenge of recovering complex physical side effects.

视频编辑物理一致性基准测试去物

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。