arXiv:2510.13271cs.CL2025-10ACL被引 1

用猜词游戏测试大模型的推理能力,发现其表现远不如人类。

Do You Get the Hint? Benchmarking LLMs on the Board Game Concept

  • 设计简单猜词游戏,考察模型的推测性推理能力。
  • 顶尖模型成功率不足40%,远低于人类90%以上。
  • 在多语言环境下,低资源语言性能进一步下降。

大型语言模型(LLMs)在多项基准测试中取得显著成果,但近期研究仍揭示其根本缺陷。本文提出Concept——一种简单的猜词类桌游,用于探测模型的溯因推理能力。结果显示,该游戏对人类而言极易完成(成功率超90%),但对当前最先进的LLMs而言仍极具挑战性(无模型超过40%成功率)。具体表现为:模型难以理解其他玩家的战略意图,且在接收到序列信息更新后无法有效修正初始假设。此外,我们扩展了跨语言评估,发现模型在荷兰语、法语和西班牙语等低资源语言中的表现相较英语进一步下降。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved striking successes on many benchmarks, yet recent studies continue to expose fundamental weaknesses. In this paper, we introduce Concept, a simple word-guessing board game, as a benchmark for probing abductive reasoning. Our results show that this game, easily solved by humans (with a success rate of over 90\%), is still very challenging for state-of-the-art LLMs (no model exceeds 40\% success rate). Specifically, we observe that LLMs struggle with interpreting other players' strategic intents, and with correcting initial hypotheses given sequential information updates. In addition, we extend the evaluation across multiple languages, and find that the LLM performance drops further in lower-resource languages (Dutch, French, and Spanish) compared to English.

大模型评测溯因推理多语言游戏基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。