arXiv:2509.09867cs.AIcs.CL2025-09ICML

让大模型当队友帮人赢牌,测试其协作能力。

LLMs as Agentic Cooperative Players in Multiplayer UNO

  • 用提示词让大模型在UNO游戏中作为合作玩家
  • 70亿参数模型仍难有效辅助他人获胜
  • 小模型也能比随机策略好,但协作效果有限

大语言模型不仅可回答问题,还能在任务中提供有用指导。但这种协助能到什么程度?我们通过让大模型参与回合制卡牌游戏UNO来检验:要求模型不追求自己赢,而是帮助另一名玩家获胜。我们构建了一个工具,使解码器类大模型可在RLCard环境中作为智能体参与游戏。模型接收完整游戏状态信息,通过简单文本提示响应,采用两种不同提示策略。评估了从10亿到700亿参数的多种模型,探索模型规模对表现的影响。结果显示,所有模型均能显著优于随机基线,但仅有少数能有效辅助他人达成目标。

原文摘要 · Abstract (English)

LLMs promise to assist humans -- not just by answering questions, but by offering useful guidance across a wide range of tasks. But how far does that assistance go? Can a large language model based agent actually help someone accomplish their goal as an active participant? We test this question by engaging an LLM in UNO, a turn-based card game, asking it not to win but instead help another player to do so. We built a tool that allows decoder-only LLMs to participate as agents within the RLCard game environment. These models receive full game-state information and respond using simple text prompts under two distinct prompting strategies. We evaluate models ranging from small (1B parameters) to large (70B parameters) and explore how model scale impacts performance. We find that while all models were able to successfully outperform a random baseline when playing UNO, few were able to significantly aid another player.

多智能体游戏AI协作大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。