arXiv:2603.03657cs.CVcs.AI2026-03被引 1

首个评测图像编辑逻辑路径的基准,揭示模型在多步推理上的短板。

InEdit-Bench: Benchmarking Intermediate Logical Pathways for Intelligent Image Editing Models

  • 构建四类任务的精细标注测试集,覆盖状态变化与动态过程。
  • 14个主流模型在路径连贯性与约束遵循上普遍表现不佳。
  • 适合关注视觉推理与智能生成模型研究者使用。

多模态生成模型在静态图像编辑任务中表现优异,但难以应对需动态推理的复杂场景,无法有效建模从初始状态到最终状态的连贯中间逻辑路径。这一能力对实现更深层次的程序性与因果理解至关重要。为此,我们提出InEdit-Bench,首个专注于图像编辑中间逻辑路径推理的评估基准。该基准包含四类基础任务:状态转换、动态过程、时间序列和科学模拟,均配有精心标注的测试用例。为支持细粒度评估,我们设计了评估标准,用于衡量生成路径的逻辑连贯性、视觉自然度及对指定路径约束的忠实度。对14个代表性图像编辑模型的全面评测显示,这些模型在该领域普遍存在显著缺陷。通过提供标准化且具有挑战性的基准,我们期望推动研究向更具动态性、推理感知和智能性的多模态生成模型发展。

原文摘要 · Abstract (English)

Multimodal generative models have made significant strides in image editing, demonstrating impressive performance on a variety of static tasks. However, their proficiency typically does not extend to complex scenarios requiring dynamic reasoning, leaving them ill-equipped to model the coherent, intermediate logical pathways that constitute a multi-step evolution from an initial state to a final one. This capacity is crucial for unlocking a deeper level of procedural and causal understanding in visual manipulation. To systematically measure this critical limitation, we introduce InEdit-Bench, the first evaluation benchmark dedicated to reasoning over intermediate pathways in image editing. InEdit-Bench comprises meticulously annotated test cases covering four fundamental task categories: state transition, dynamic process, temporal sequence, and scientific simulation. Additionally, to enable fine-grained evaluation, we propose a set of assessment criteria to evaluate the logical coherence and visual naturalness of the generated pathways, as well as the model's fidelity to specified path constraints. Our comprehensive evaluation of 14 representative image editing models on InEdit-Bench reveals significant and widespread shortcomings in this domain. By providing a standardized and challenging benchmark, we aim for InEdit-Bench to catalyze research and steer development towards more dynamic, reason-aware, and intelligent multimodal generative models.

图像编辑逻辑推理基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。