让大模型通过语言博弈学习策略决策,胜率超GPT-4o。
Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game
- 基于语言游戏设计多智能体交互框架,直接优化策略与语言统一能力。
- 9人狼人杀中平均胜率达61%,较GPT-4o提升23%。
- 表现接近人类,盲测检测率仅49%,适合研究智能体协作与博弈。
实现通用人工智能(AGI)需要具备战略决策与灵活沟通能力的智能体。受维特根斯坦《哲学研究》中语言游戏理论启发,我们提出语言智能体可通过上下文交互学习,而非传统分阶段决策与语言表达分离的框架。以考验语言理解、策略互动与适应性的社交推理游戏狼人杀为场景,我们构建了多智能体卡尼曼与特沃斯基优化(MaKTO)方法。MaKTO在大量9人狼人杀对局中生成无配对的优质与劣质响应,再利用KTO优化模型决策过程。实验表明,在9人狼人杀游戏中,MaKTO跨多种模型平均胜率达61%,相较GPT-4o提升23.0%,优于两阶段强化学习代理10.9%。值得注意的是,其表现接近人类水平,对专家玩家胜率达60%,且在图灵式盲测中仅被识别出49%。
原文摘要 · Abstract (English)
Achieving Artificial General Intelligence (AGI) requires AI agents that can not only make stratigic decisions but also engage in flexible and meaningful communication. Inspired by Wittgenstein's language game theory in Philosophical Investigations, we propose that language agents can learn through in-context interaction rather than traditional multi-stage frameworks that separate decision-making from language expression. Using Werewolf, a social deduction game that tests language understanding, strategic interaction, and adaptability, we develop the Multi-agent Kahneman & Tversky's Optimization (MaKTO). MaKTO engages diverse models in extensive gameplay to generate unpaired desirable and unacceptable responses, then employs KTO to refine the model's decision-making process. In 9-player Werewolf games, MaKTO achieves a 61% average win rate across various models, outperforming GPT-4o and two-stage RL agents by relative improvements of 23.0% and 10.9%, respectively. Notably, MaKTO also demonstrates human-like performance, winning 60% against expert players and showing only 49% detectability in Turing-style blind tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。