用文字版猜谜游戏测试大模型推理能力,发现表现普遍不佳。
How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
- 构建文字版《妙探闯通关》游戏环境,评估多步逻辑推理。
- 18场模拟游戏中仅4次正确破案,显示持续推理能力弱。
- 微调逻辑题未提升表现,反而增加错误推理量。
大语言模型在推断‘谁干的’这类问题上表现困难。本文构建了一个基于规则的文字版多智能体《妙探闯通关》游戏,作为评估多步逻辑推理能力的测试平台,共使用GPT-4o-mini和Gemini-2.5-Flash六种代理进行测试。我们进一步探究在结构化逻辑谜题上微调是否能提升游戏中的推理与表现。在18场模拟游戏中,仅有4次成功破案,表明模型难以在整个游戏过程中保持一致的推理能力。此外,微调并未可靠提升性能,甚至在某些情况下增加了推理量但未提高准确性。
原文摘要 · Abstract (English)
Deducing whodunit proves challenging for LLM agents. In this paper, we implement a text-based multi-agent version of the classic board game Clue as a rule-based testbed for evaluating multi-step deductive reasoning, with six agents drawn from GPT-4o-mini and Gemini-2.5-Flash. We further investigate whether fine-tuning on structured logic puzzles transfers to improved in-game reasoning and gameplay. Across 18 simulated games, agents achieve only four correct wins, indicating difficulty in maintaining consistent deductive reasoning over the course of a full game. Additionally, we find that fine-tuning does not reliably improve performance and, in some cases, appears to increase reasoning volume without improving reasoning precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。