提出新评估框架,让机器判断图像视频去物更像人眼。
PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media

- 用滑窗比特征和帧间分布追踪,衡量去物区域的视觉连贯性。
- 在80个带运动增强视频上测试,相比旧方法更贴近人眼判断。
- 适合研究去物算法的开发者和评测人员使用。
图像与视频中的物体移除评估仍具挑战性,因该任务本质为一对多,而现有度量常与人类感知不符。全参考度量偏向复制粘贴行为而非真实擦除;无参考度量存在系统性偏差,如偏好模糊结果;全局时序度量对编辑区域内的局部伪影不敏感。为此,我们提出RC(Removal Coherence)双指标:RC-S通过掩码区与背景区的滑窗特征对比衡量空间连贯性,RC-T通过相邻帧共享恢复区域内的分布追踪衡量时间一致性。为验证RC并支持社区基准测试,我们进一步构建了两层真实世界基准PROVE-Bench:PROVE-M包含80个经运动增强的配对视频,PROVE-H为100个无真值的高难度子集。RC指标与PROVE-Bench共同构成视觉媒体去物评估框架PROVE。跨多样图像与视频基准的实验表明,RC在与人类判断的一致性上显著优于现有评估协议。
原文摘要 · Abstract (English)
Evaluating object removal in images and videos remains challenging because the task is inherently one-to-many, yet existing metrics frequently disagree with human perception. Full-reference metrics reward copy-paste behaviors over genuine erasure; no-reference metrics suffer from systematic biases such as favoring blurry results; and global temporal metrics are insensitive to localized artifacts within edited regions. To address these limitations, we propose RC (Removal Coherence), a pair of perception-aligned metrics: RC-S, which measures spatial coherence via sliding-window feature comparison between masked and background regions, and RC-T, which measures temporal consistency via distribution tracking within shared restored regions across adjacent frames. To validate RC and support community benchmarking, we further introduce PROVE-Bench, a two-tier real-world benchmark comprising PROVE-M, an 80-video paired dataset with motion augmentation, and PROVE-H, a 100-video challenging subset without ground truth. Together, RC metrics and PROVE-Bench form the PROVE (Perceptual RemOVal cohErence) evaluation framework for visual media. Experiments across diverse image and video benchmarks demonstrate that RC achieves substantially stronger alignment with human judgments than existing evaluation protocols. Project page: https://xiaomi-research.github.io/prove/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。