arXiv:2510.17681cs.CVcs.AI2025-10被引 13

新基准测试揭示图像编辑仍缺乏物理真实感

PICABench: How Far Are We from Physically Realistic Image Editing?

  • 构建八维度物理真实性评测体系,覆盖光学、力学等场景
  • 主流模型在移除物体时仍无法正确处理阴影与反射
  • 提出视频学习物理规律的解决方案,提供百万级训练数据

近年来图像编辑取得显著进展,现代模型已能执行复杂指令。然而,除了完成编辑任务外,物理效应才是生成真实感的关键。例如删除物体应同时消除其阴影、反射及与周围物体的交互。现有模型与基准主要关注指令完成度,忽视了这些物理效应。为此,我们提出PICABench,系统评估常见编辑操作(添加、移除、属性修改等)在光学、力学和状态变化等八个子维度上的物理真实性。我们进一步提出PICAEval评估协议,采用视觉语言模型作为评判者,结合逐例、区域级人工标注与提问。除基准构建外,我们通过视频学习物理规律,构建了包含10万张图像的PICA-100K训练数据集。评估主流模型后发现,物理真实性仍是重大挑战。我们希望该基准与方案能推动研究从简单内容编辑迈向物理一致性生成。

原文摘要 · Abstract (English)

Image editing has achieved remarkable progress recently. Modern editing models could already follow complex instructions to manipulate the original content. However, beyond completing the editing instructions, the accompanying physical effects are the key to the generation realism. For example, removing an object should also remove its shadow, reflections, and interactions with nearby objects. Unfortunately, existing models and benchmarks mainly focus on instruction completion but overlook these physical effects. So, at this moment, how far are we from physically realistic image editing? To answer this, we introduce PICABench, which systematically evaluates physical realism across eight sub-dimension (spanning optics, mechanics, and state transitions) for most of the common editing operations (add, remove, attribute change, etc.). We further propose the PICAEval, a reliable evaluation protocol that uses VLM-as-a-judge with per-case, region-level human annotations and questions. Beyond benchmarking, we also explore effective solutions by learning physics from videos and construct a training dataset PICA-100K. After evaluating most of the mainstream models, we observe that physical realism remains a challenging problem with large rooms to explore. We hope that our benchmark and proposed solutions can serve as a foundation for future work moving from naive content editing toward physically consistent realism.

图像编辑物理真实性基准测试视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。