测试图像生成模型解决视觉谜题的能力,发现主流模型表现有限。
GenPuzzle: Benchmarking Visual Reasoning in Image Generation Models

- 设计12类视觉谜题任务,要求生成图像完整还原输入并逻辑求解。
- 三款顶尖生成模型最高仅40.57分,暴露出逻辑与几何错误频发。
- 用多模态大模型和人工评分验证,适合评估视觉推理能力的模型。
近期图像生成系统融合了多模态理解、推理与合成能力,可能已超越单纯渲染合理场景。但现有评估多关注美学、提示对齐、组合性或文本回答,难以判断其是否真正具备视觉问题求解能力。我们提出GenPuzzle,一个以推理为核心的图像生成基准,包含2,005个跨12类任务的视觉谜题,涵盖模式补全、空间构建、迷宫、数独、非欧诺格拉姆、七巧板、棋盘游戏、火柴棒谜题、正投影及数学视觉证明等。每项任务需在保留输入状态的前提下,生成符合逻辑的有效解决方案图像。采用任务特异的评估协议:离散网格输出由程序转录验证,复杂视觉输出则使用分级、多维或二元的多模态大语言模型(MLLM)评分标准。通过与人类参考评分的一致性筛选自动评判器。在三款前沿生成模型中,最强模型仅达40.57分的宏观平均成绩,暴露出逻辑、几何、状态保持与指令执行方面的频繁失败。GenPuzzle为衡量图像生成从绘图向视觉求解演进提供了测试平台。
原文摘要 · Abstract (English)
Recent image generation systems increasingly combine multimodal understanding, reasoning, and synthesis, suggesting that they may do more than render plausible scenes. Yet existing evaluations emphasize aesthetics, prompt alignment, compositionality, or text-based answers, leaving unclear whether these systems can solve visual problems and faithfully express solutions in pixels. We introduce GenPuzzle, a benchmark for reasoning-centric image generation. GenPuzzle contains 2,005 problems across 12 tracks, spanning pattern completion, spatial construction, mazes, Sudoku, nonograms, tangrams, board games, matchstick puzzles, orthographic projection, and mathematical visual proof. Each task provides a visual puzzle and requires an image output that preserves the input state while executing a logically valid solution. GenPuzzle uses task-specific evaluation protocols: discrete grid outputs are transcribed and verified programmatically, while visually complex outputs are assessed with tiered, multidimensional, or binary multimodal large language model (MLLM) rubrics. We further select the automatic judge by measuring agreement with human reference scores. Across three frontier generators, the strongest model reaches only 40.57 Macro Overall, revealing frequent failures in logic, geometry, state preservation, and instruction execution. GenPuzzle provides a testbed for measuring progress from image rendering toward visual problem solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。