arXiv:2604.22868cs.CVcs.AI2026-04中稿 · ICLR被引 1

提出单步图像编辑新范式,揭示模型视觉规划能力短板

Probing Visual Planning in Image Editing Models

论文配图:Probing Visual Planning in Image Editing Models
图 1 · 摘自论文原文
  • 将视觉规划重构为单步图像变换,避免逐帧生成耗时
  • 在抽象迷宫与后棋问题上,微调后模型跨尺度泛化能力显著提升
  • 首次实现对生成模型逻辑正确性与像素精度的自动评估

视觉规划是人类智能的关键组成部分,尤其在需要复杂空间推理的任务中。然而,机器学习中该问题常被以语言为中心的方法处理。尽管近期研究展示了纯视觉方法的潜力,但其基于逐步生成的规划范式存在严重计算效率问题。本文提出EAR——一种将图像编辑视为推理的新范式,将视觉规划重构为单步图像转换。为分离内在推理与视觉识别,我们采用抽象谜题作为探测任务,并引入AMAZE数据集,包含经典迷宫和后棋问题,涵盖不同且互补的视觉规划形式。AMAZE的抽象特性支持对自回归与扩散模型在像素级保真度和逻辑有效性上的自动评估。我们评估了领先的专有与开源编辑模型。结果表明,所有模型在零样本设置下表现不佳;在基础规模上微调后,模型在同域及跨域尺度、几何结构上均展现出显著泛化能力。然而,即使在高端硬件上运行的最佳模型,仍无法达到人类求解者的零样本效率,凸显神经视觉推理中的持续差距。

原文摘要 · Abstract (English)

Visual planning represents a crucial facet of human intelligence, especially in tasks that require complex spatial reasoning and navigation. Yet, in machine learning, this inherently visual problem is often tackled through a verbal-centric lens. While recent research demonstrates the promise of fully visual approaches, they suffer from significant computational inefficiency due to the step-by-step planning-by-generation paradigm. In this work, we present EAR, an editing-as-reasoning paradigm that reformulates visual planning as a single-step image transformation. To isolate intrinsic reasoning from visual recognition, we employ abstract puzzles as probing tasks and introduce AMAZE, a procedurally generated dataset that features the classical Maze and Queen problems, covering distinct, complementary forms of visual planning. The abstract nature of AMAZE also facilitates automatic evaluation of autoregressive and diffusion-based models in terms of both pixel-wise fidelity and logical validity. We assess leading proprietary and open-source editing models. The results show that they all struggle in the zero-shot setting, finetuning on basic scales enables remarkable generalization to larger in-domain scales and out-of-domain scales and geometries. However, our best model that runs on high-end hardware fails to match the zero-shot efficiency of human solvers, highlighting a persistent gap in neural visual reasoning.

视觉规划图像编辑生成模型自动评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。