arXiv:2601.16276cs.CLcs.AI2026-01

让大模型通过多轮对话学会战略性决策,提升协作与谈判能力。

GameTalk: Training LLMs for Strategic Conversation

  • 用多轮对话的全局奖励训练模型,优化长期战略目标。
  • 在复杂博弈任务中表现显著优于未训练模型,尤其配合奖励设计时效果更佳。
  • 适合研究对话智能、多智能体协作与策略推理的开发者和研究人员。

在多智能体环境中进行战略决策是大型语言模型(LLMs)面临的核心挑战,尤其是在需要长时间对话来实现协调与谈判的情况下。尽管近期工作已探索了LLM在孤立决策任务中的应用,但针对通过对话实现长期目标优化的研究仍较少。本文提出「GameTalk」框架,旨在通过多轮交互训练LLM做出战略性决策。不同于以往聚焦单轮目标或静态动作预测的方法,我们采用GRPO、DPO和STaR等微调方法,引入依赖于完整交互过程的奖励信号,以优化全局目标。我们在一系列逐步复杂的游戏中评估该方法,这些游戏旨在考验推理、协作与对手建模能力。结果表明,GameTalk显著优于未经训练的模型,尤其在奖励塑形条件下,其中DPO始终带来最佳提升。这些发现表明,对话式微调是推动LLM在互动环境中进行推理、协商与行动的有前景路径。

原文摘要 · Abstract (English)

Strategic decision-making in multi-agent settings is a key challenge for large language models (LLMs), particularly when coordination and negotiation must unfold over extended conversations. While recent work has explored the use of LLMs in isolated decision tasks, little attention has been given to optimizing long-term objectives through dialogue. We introduce \textbf{GameTalk}, a framework for training LLMs to make strategic decisions via multi-turn interactions. Unlike prior work that focuses on single-turn objectives or static action prediction, we train LLMs to optimize a global objective across full conversations. We achieve this by adapting fine-tuning methods like GRPO, DPO, and STaR to incorporate reward signals that depend on the entire interaction. We evaluate this approach on a suite of increasingly complex games, designed to stress different aspects of reasoning, coordination, and opponent modeling. Our results show that GameTalk significantly outperforms untrained models, especially under reward shaping, with DPO consistently yielding the strongest gains. These findings position conversational fine-tuning as a promising path for LLMs to reason, negotiate, and act in interactive environments.

战略对话多智能体对话微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。