arXiv:2506.06211cs.CLcs.AI2025-06被引 4

构建首个多模态开放推理基准,评估模型解谜能力。

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts

  • 设计667个开放式解谜题,含多模态线索与逐步推理轨迹。
  • 顶尖模型仅18%正确率,步级准确率40%,远低于人类高手。
  • 标注推理过程可提升小模型表现,助力视觉推理研究。

Puzzlehunts 是一类复杂、多步骤且无明确问题定义的谜题。与传统有明确指令和受限环境的任务不同,Puzzlehunts 要求从多模态证据中发现潜在问题结构并进行迭代推理,类似科学发现、探索性数据分析或调查式问题解决。尽管基础模型取得进展,其在开放设定下的表现仍缺乏验证。我们提出 PuzzleWorld,一个包含667个解谜风格题目的综合性基准,用于评估逐步、开放和创造性的多模态推理能力。每个题目配有最终答案、详细推理轨迹和认知技能标签,支持全面评测与细粒度诊断分析。当前最先进模型的最终答案准确率仅为1-4%。在 PuzzleWorld 上,最佳模型仅能解决18%的谜题,步级准确率为40%,相当于人类新手水平,显著落后于解谜高手。通过分析推理标注,我们发现微调小模型可将步级准确率从4%提升至11%,并在下游视觉推理任务中带来性能提升。详细错误分析显示,当前模型存在短视推理、受限于语言推理能力,且缺乏对视觉与空间推理至关重要的草图生成能力。我们已在 https://github.com/MIT-MI/PuzzleWorld 开源 PuzzleWorld,以推动更通用、开放和创造性推理系统的发展。

原文摘要 · Abstract (English)

Puzzlehunts are a genre of complex, multi-step puzzles lacking well-defined problem definitions. In contrast to conventional reasoning benchmarks consisting of tasks with clear instructions and constrained environments, puzzlehunts requires discovering the underlying problem structure from multimodal evidence and iterative reasoning, mirroring real-world domains such as scientific discovery, exploratory data analysis, or investigative problem-solving. Despite progress in foundation models, their performance on open-ended settings remains largely untested. We introduce PuzzleWorld, a comprehensive benchmark of 667 puzzlehunt-style problems designed to assess step-by-step, open-ended, and creative multimodal reasoning. Each puzzle is annotated with the final solution, detailed reasoning traces, and cognitive skill labels, enabling holistic benchmarking and fine-grained diagnostic analysis. Most state-of-the-art models achieve only 1-4% final answer accuracy. On PuzzleWorld, the best model solves only 18% of puzzles and reaches 40% stepwise accuracy, matching human puzzle novices but falling significantly behind puzzle enthusiasts. To demonstrate the value of our reasoning annotations, we show that fine-tuning a small model on reasoning traces boosts stepwise accuracy from 4% to 11%, which translates to improvements in downstream visual reasoning tasks. Our detailed error analysis reveals that current models exhibit myopic reasoning, are bottlenecked by the limitations of language-based inference, and lack sketching capabilities crucial for visual and spatial reasoning. We release PuzzleWorld at https://github.com/MIT-MI/PuzzleWorld to support future work on building more general, open-ended, and creative reasoning systems.

多模态开放推理解谜基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。