arXiv:2505.15712cs.CL2025-05EMNLP被引 2

用侦探游戏评测大模型的逻辑推理能力,发现现有方法仍有局限。

TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games

  • 基于侦探游戏设计推理任务,要求模型从长文本中找证词与证据的矛盾。
  • 12个主流大模型在任务中表现不佳,尤其在长上下文和多步骤推理时退化明显。
  • 适合研究逻辑推理、可信AI或游戏化评测的学者和开发者。

本文提出TurnaboutLLM,一个基于侦探游戏《逆转裁判》和《弹丸论破》互动玩法的框架与数据集,用于评估大语言模型(LLMs)的演绎推理能力。该任务要求模型在长篇叙事上下文中识别证词与证据之间的矛盾,挑战性源于答案空间庞大及推理类型多样。我们在该数据集上评估了12个前沿大模型,结果表明,现有提升推理能力的常见策略(如深度思考、思维链提示)效果有限。同时,实验揭示了上下文长度、推理步骤数和答案空间大小对模型表现的影响存在差异。总体而言,TurnaboutLLM为大模型在复杂叙事环境中的演绎推理能力提出了实质性挑战。

原文摘要 · Abstract (English)

This paper introduces TurnaboutLLM, a novel framework and dataset for evaluating the deductive reasoning abilities of Large Language Models (LLMs) by leveraging the interactive gameplay of detective games Ace Attorney and Danganronpa. The framework tasks LLMs with identifying contradictions between testimonies and evidences within long narrative contexts, a challenging task due to the large answer space and diverse reasoning types presented by its questions. We evaluate twelve state-of-the-art LLMs on the dataset, hinting at limitations of popular strategies for enhancing deductive reasoning such as extensive thinking and Chain-of-Thought prompting. The results also suggest varying effects of context size, the number of reasoning step and answer space size on model performance. Overall, TurnaboutLLM presents a substantial challenge for LLMs' deductive reasoning abilities in complex, narrative-rich environments.

逻辑推理大模型评测游戏化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。