arXiv:2412.11373cs.AIcs.CL2024-12被引 10

用猜词游戏测试大模型的推理能力,发现不同模型表现各异且合作更灵活。

Codenames as a Benchmark for Large Language Models

  • 用经典猜词游戏设计评测任务,考察语言理解与共情推理。
  • GPT-4o等模型在多种布局下表现不一,展现角色适应性差异。
  • 多模型协作比传统方法更泛化,适合复杂团队场景。

本文提出将流行的词语类桌游Codenames作为评估大语言模型(LLMs)推理能力的基准。该游戏对人工智能提出了严峻挑战,要求具备复杂的语言理解、心理理论和认知推理能力。此前的智能体主要依赖词嵌入技术,词汇范围有限且泛化能力差。尽管LLMs在语言任务中表现出更强的推理与理解能力,但在横向思维方面仍有不足。我们评估了GPT-4o、Gemini 1.5、Claude 3.5 Sonnet和Llama 3.1等主流模型在多种棋盘设置下的表现。结果表明,虽部分模型整体表现更优,但各模型在游戏过程中展现出不同的涌现行为,并在特定角色上表现出色。此外,我们还测试了多模型协同合作的表现,证明LLM智能体比以往方法更能适应多样队友,具有更强的可迁移性。

原文摘要 · Abstract (English)

In this paper, we propose the use of the popular word-based board game Codenames as a suitable benchmark for evaluating the reasoning capabilities of Large Language Models (LLMs). Codenames presents a highly interesting challenge for achieving successful AI performance, requiring both a sophisticated understanding of language, theory of mind, and epistemic reasoning capabilities. Prior attempts to develop agents for Codenames have largely relied on word embedding techniques, which have a limited vocabulary range and perform poorly when paired with differing approaches. LLMs have demonstrated enhanced reasoning and comprehension capabilities for language-based tasks, but can still suffer in lateral thinking challenges. We evaluate the capabilities of several state-of-the-art LLMs, including GPT-4o, Gemini 1.5, Claude 3.5 Sonnet, and Llama 3.1, across a variety of board setups. Our results indicate that while certain LLMs perform better than others overall, different models exhibit varying emergent behaviours during gameplay and excel at specific roles. We also evaluate the performance of different combinations of LLMs when playing cooperatively together, demonstrating that LLM agents are more generalisable to a wider range of teammates than prior techniques.

大模型评测推理能力游戏基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。