用新游戏BrainKing评估大模型在信息不全时的推理能力。
Do Large Language Models have Problem-Solving Capability under Incomplete Information Scenarios?
- 设计新游戏BrainKing融合猜谜与反误导机制
- 大模型在高难度下正确率仅41%,暴露推理短板
- 适合研究大模型认知局限与交互式推理的学者
大语言模型在信息不全场景下的问题求解能力评估日益重要,涉及提问、知识检索、错误检测和路径规划等能力。现有研究多聚焦于类似‘二十个问题’的游戏,但这些游戏无需识别误导性线索,而真实场景中此类能力至关重要。此外,如‘谁是卧底’类游戏主观性强,难以客观评估。为此,本文提出基于‘谁是卧底’与‘二十个问题’的新游戏BrainKing,用于评估大模型在有限是非问答和潜在误导回答下的目标实体识别能力。通过设置简单、中等、困难三种难度模式,全面评估大模型表现。结果揭示了大模型在BrainKing中的能力边界,为理解其问题求解水平提供了重要洞见。
原文摘要 · Abstract (English)
The evaluation of the problem-solving capability under incomplete information scenarios of Large Language Models (LLMs) is increasingly important, encompassing capabilities such as questioning, knowledge search, error detection, and path planning. Current research mainly focus on LLMs' problem-solving capability such as ``Twenty Questions''. However, these kinds of games do not require recognizing misleading cues which are necessary in the incomplete information scenario. Moreover, the existing game such as ``Who is undercover'' are highly subjective, making it challenging for evaluation. Therefore, in this paper, we introduce a novel game named BrainKing based on the ``Who is undercover'' and ``Twenty Questions'' for evaluating LLM capabilities under incomplete information scenarios. It requires LLMs to identify target entities with limited yes-or-no questions and potential misleading answers. By setting up easy, medium, and hard difficulty modes, we comprehensively assess the performance of LLMs across various aspects. Our results reveal the capabilities and limitations of LLMs in BrainKing, providing significant insights of LLM problem-solving levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。