arXiv:2503.15235cs.CLcs.AI2025-03被引 4

用大模型玩猜间谍游戏,不训练也能高分

Exploring Large Language Models for Word Games:Who is the Spy?

  • 用思维链调度框架让大模型推理角色和伪装身份
  • 多数据集测试显示模型表现显著提升
  • 适合研究情境推理与社交智能的学者参考

词语游戏在自然语言处理、博弈论等相关领域具有重要研究价值,因其规则明确且情境性强。本研究探索大语言模型(LLM)在词语游戏中的有效应用,提出一种无需训练的框架。以经典游戏「谁是卧底」(Who is the Spy)为例,引入基于思维链(CoT)的调度机制,使LLM在推断角色词和伪装身份等任务中表现优异。通过游戏成功率和分析结果准确率评估框架性能。实验结果验证了该框架的有效性,在多个数据集上均显著提升LLM表现。工作凸显了大模型在结构化游戏环境中掌握情境推理与社交互动的潜力。代码已开源:https://github.com/ct-wei/Who-is-The-Spy。

原文摘要 · Abstract (English)

Word games hold significant research value for natural language processing (NLP), game theory, and related fields due to their rule-based and situational nature. This study explores how large language models (LLMs) can be effectively involved in word games and proposes a training-free framework. "Shei Shi Wo Di" or "Who is the Spy" in English, is a classic word game. Using this game as an example, we introduce a Chain-of-Thought (CoT)-based scheduling framework to enable LLMs to achieve excellent performance in tasks such as inferring role words and disguising their identities. We evaluate the framework's performance based on game success rates and the accuracy of the LLM agents' analytical results. Experimental results affirm the framework's effectiveness, demonstrating notable improvements in LLM performance across multiple datasets. This work highlights the potential of LLMs in mastering situational reasoning and social interactions within structured game environments. Our code is publicly available at https://github.com/ct-wei/Who-is-The-Spy.

大模型推理游戏思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。