arXiv:2606.00103cs.AI2026-06

用可执行游戏评估大模型的动态推理能力,看它如何边问边学、边改边答。

Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games

  • 让大模型通过提问与环境互动,逐步收集证据并更新判断。
  • 474个游戏测试显示,模型在成功率和交互效率上差异显著。
  • 模型对错误假设的修正能力弱,适合研究推理鲁棒性的学者。

我们提出一种多轮交互式推理评估框架,将推理视为主动获取证据与更新信念的过程。大模型仅知任务规则,需向隐藏环境发出针对性提问,随时间整合部分观测结果,并决定何时提交最终答案。除标准成功率与交互效率外,还评估在受控上下文扰动下的情境鲁棒性,以及通过反事实修订和必要性判断体现的元认知适应能力。该框架被实例化为包含474个可执行游戏的基准,每类游戏在五个固定配置搜索空间下评估,对应五种难度等级,并测试了多种前沿大模型。结果表明,该基准具有高度区分性,不仅暴露了模型在成功率上的差异,也揭示了交互效率的显著差距。此外,实验显示上下文扰动导致中等但一致的性能下降,而反事实修订与必要性判断则引发更大程度的性能损失。

原文摘要 · Abstract (English)

We introduce a multi-turn interactive framework for reasoning evaluation that treats reasoning as active evidence acquisition and belief updating. Wherein, LLMs receive only the task rules, must issue targeted queries to a hidden environment, integrate partial observations over time, and decide when to submit a final answer. Beyond standard success rate and interaction efficiency, we evaluate contextual robustness under controlled contextual perturbations, and metacognitive adaptation through counterfactual revision and necessity judgment. We instantiate the framework as a benchmark of 474 executable games, each evaluated under five fixed configuration search spaces corresponding to five difficulty levels, and evaluate a broad set of frontier LLMs. Results show that the benchmark is highly discriminative, exposing large differences not only in success rate but also in interaction efficiency. Moreover, we empirically show that contextual perturbations cause moderate but consistent declines, whereas counterfactual revision and necessity judgment lead to much larger drops.

大模型推理交互评估元认知游戏基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。