构建真实物理场景图像编辑评测基准,揭示模型短板并提出新解法
PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

- 基于分层分类体系设计238个真实视频提取实例+35个反物理合成例
- 现有顶尖模型在物理推理上表现差,仅平均正确率不足60%
- 无需训练的PhyWorld利用视频生成过程实现推理,性能超越对比方法
尽管基于指令的图像编辑已取得显著进展,但现有评测基准缺乏对物理推理能力的全面评估,而这正是处理真实世界场景的关键。为此,我们提出PhyEditBench,一个用于评估编辑模型物理理解能力的多阶段基准。依据层级分类体系,建立4个主类与12个子类,包含238个从视频中精心提取的高质量、高分辨率真实世界实例,以及35个合成的反物理实例。对当前最先进编辑方法的实证分析显示其在物理推理方面存在显著局限。我们进一步提出一种无需训练的基线方法PhyWorld,采用测试时缩放和潜在空间降维策略。PhyWorld在多项指标上优于可比模型,表明视频生成过程可有效充当图像编辑的推理机制。项目主页见https://github.com/Previsior/PhyEditBench。
原文摘要 · Abstract (English)
While instruction-based image editing, enabled by multi-modal generative models, has advanced significantly, existing benchmarks lack a comprehensive evaluation of physics-based reasoning, a critical capability for handling real-world scenarios. To address this, we introduce PhyEditBench, a benchmark designed to assess the physical understanding of editing models. Guided by a hierarchical taxonomy, we establish 4 primary classes and 12 subclasses. It comprises 238 high-quality, high-resolution, real-world instances meticulously extracted from videos to capture authentic physical dynamics, alongside 35 synthetic Anti-Physics instances. Our empirical analysis of current SOTA editing methods exposes substantial limitations in their physics-based reasoning. We further propose a training-free baseline named PhyWorld that uses test-time scaling and a latent reduction strategy. PhyWorld outperforms comparable models and suggests that the video generation process can effectively serve as a reasoning mechanism for image editing. The project page is available at https://github.com/Previsior/PhyEditBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。