用迷题测试大模型的主动推理能力,发现其远不如人类。
What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
- 设计互动谜题框架,动态评估模型假设构建与修正能力。
- 800个双语谜题中,顶尖模型表现显著落后于人类。
- 提出多维评测协议,可精准捕捉推理缺陷与盲点。
我们研究大型语言模型在想象力推理方面的能力——即在信息稀疏环境下主动构建、检验和修正假设。现有基准多为静态或聚焦社交推理,无法捕捉这一过程的动态探索特性。为此,我们基于经典「乌龟汤」游戏,构建包含基准、智能体与评估协议的完整研究框架。提出TurtleSoup-Bench,首个大规模、双语、交互式想象力推理基准,包含800个来自网络与专家作者的谜题。同时设计Mosaic-Agent,用于评估模型在此场景下的表现。开发多维度评估协议,衡量逻辑一致性、细节补全与结论匹配度。对主流LLM的实验揭示其明显能力局限、共性失败模式,以及与人类间的显著差距。本工作为理解大模型的探索性推理提供了新视角,并奠定未来研究基础。
原文摘要 · Abstract (English)
We investigate the capacity of Large Language Models (LLMs) for imaginative reasoning--the proactive construction, testing, and revision of hypotheses in information-sparse environments. Existing benchmarks, often static or focused on social deduction, fail to capture the dynamic, exploratory nature of this reasoning process. To address this gap, we introduce a comprehensive research framework based on the classic "Turtle Soup" game, integrating a benchmark, an agent, and an evaluation protocol. We present TurtleSoup-Bench, the first large-scale, bilingual, interactive benchmark for imaginative reasoning, comprising 800 turtle soup puzzles sourced from both the Internet and expert authors. We also propose Mosaic-Agent, a novel agent designed to assess LLMs' performance in this setting. To evaluate reasoning quality, we develop a multi-dimensional protocol measuring logical consistency, detail completion, and conclusion alignment. Experiments with leading LLMs reveal clear capability limits, common failure patterns, and a significant performance gap compared to humans. Our work offers new insights into LLMs' imaginative reasoning and establishes a foundation for future research on exploratory agent behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。