arXiv:2504.02826cs.CV2025-04NeurIPS被引 93

首个评估视觉编辑中推理能力的基准,揭示当前模型在复杂指令下表现极弱。

Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing

  • 构建包含时序、因果、空间、逻辑四类推理的视觉编辑评测集
  • 最强模型GPT-4o-Image在复杂任务上准确率仅28.8%
  • 支持人工与模型评分双验证,适合研究视觉推理与生成的学者

大型多模态模型在视觉理解与生成方面取得显著进展,但在通用视觉编辑任务中仍面临挑战,尤其在遵循复杂指令、保持外观一致性以及支持灵活输入格式方面。为探究这一差距,我们提出RISEBench,首个用于评估推理驱动视觉编辑(RISE)的基准。RISEBench聚焦于时序、因果、空间和逻辑四类关键推理能力,精心构建高质量测试用例,并提出一个稳健的评估框架,通过人工评审与‘大模型作为裁判’方法,评估指令推理、外观一致性和视觉合理性。我们对九个主流视觉编辑模型(含开源与专有模型)进行了实验。结果表明,当前模型在基于推理的编辑任务中存在显著缺陷。即使是最强大的评估模型GPT-4o-Image,准确率也仅为28.8%。RISEBench有效揭示了现有编辑模型的局限性,提供了重要洞见,并指明了未来推理感知视觉编辑的发展方向。代码与数据已公开于https://github.com/PhoenixZ810/RISEBench。

原文摘要 · Abstract (English)

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE). RISEBench focuses on four key reasoning categories: Temporal, Causal, Spatial, and Logical Reasoning. We curate high-quality test cases for each category and propose an robust evaluation framework that assesses Instruction Reasoning, Appearance Consistency, and Visual Plausibility with both human judges and the LMM-as-a-judge approach. We conducted experiments evaluating nine prominent visual editing models, comprising both open-source and proprietary models. The evaluation results demonstrate that current models face significant challenges in reasoning-based editing tasks. Even the most powerful model evaluated, GPT-4o-Image, achieves an accuracy of merely 28.8%. RISEBench effectively highlights the limitations of contemporary editing models, provides valuable insights, and indicates potential future directions for the field of reasoning-aware visual editing. Our code and data have been released at https://github.com/PhoenixZ810/RISEBench.

视觉编辑多模态推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。