测试图像编辑模型能否像人一样解图像逻辑题,发现视觉规则并自动完成任务。
Image-Space Rule Discovery

- 设计新基准WISRD,用图像指令测试模型推理与编辑能力
- 纳米香蕉Pro表现最佳,无参考任务通过率达48.7%
- 模型可依赖图像内提示,即使外部指令模糊仍能部分响应
本文以人类智力测验为灵感,探讨图像编辑模型是否能在图像空间中发现视觉规则并端到端完成问题求解。我们提出WISRD基准,包含11个核心任务和4个附加推理压力测试,覆盖局部标记、填空、复制、计数及无编辑抑制等场景。评估显示,前沿模型中纳米香蕉Pro表现最优,在共享的V0--V3无参考子集上,其自动严格代理通过率为48.7%;而Qwen-Image-Edit、FLUX.2 Klein 4B API等分别为13.4%、11.5%和11.3%,InstructPix2Pix为0.0%。分析表明,当前模型可在无外部提示或提示通用时,仍部分依赖图像内渲染指令。小规模诊断中,纳米香蕉Pro在4×4数独任务上达到70.0%准确率,在公开RAVEN模式发现任务中达22.9%。
原文摘要 · Abstract (English)
Can image-editing models discover visual rules in image space and complete problem-solving end-to-end? We tackle this question in the spirit of a human worksheet test (e.g., an IQ test), using problems that require models to read image-based instructions, recognize the problem, infer the answer, bind it to the correct destination, control output count, suppress unnecessary edits, and preserve the input and format. We introduce WISRD, a Worksheet Image-Space Rule Discovery benchmark with 11 core tasks under eight information conditions, spanning localized marking, filling, copying, counting, and no-edit suppression, together with four supplementary reasoning-stress probes for multi-step spatial manipulation, abstract pattern reasoning, logical inference, and constraint-based problem solving. We identify three key findings as follows. (i) Among the frontier image-editing models evaluated, Nano Banana Pro achieves the highest score. On the shared V0--V3 no-reference subset, the Auto-Strict proxy pass rates are 48.7% for Nano Banana Pro, 13.4\% for Qwen-Image-Edit, 11.5% for FLUX.2 Klein 4B API, 11.3% for FLUX.2 Klein 4B open-weight, and 0.0% for InstructPix2Pix. (ii) Analysis reveals that current image-editing models can partially rely on rendered in-image instructions even when the external prompt is absent or merely generic. (iii) In small supplementary diagnostics, Nano Banana Pro achieves 70.0% on 4-by-4 Sudoku and 22.9% on public RAVEN pattern-discovery items in image space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。