一个模型同时扮演对话各方,通过自对弈学会复杂社交能力。
One Model, All Roles: Multi-Turn, Multi-Agent Self-Play Reinforcement Learning for Conversational Social Intelligence
- 单模型自对弈,动态角色切换学习长期社交目标。
- 在SOTOPIA和狼人杀中实现共情、说服等自发社交行为。
- 无需人工标注,适合研究群体对话中的智能演化。
本文提出OMAR:一个单一模型在多轮多主体对话自对弈中发展社交智能的强化学习框架。与传统静态单轮优化不同,该框架使同一模型可同时扮演对话各方,直接从动态互动中学习长期目标与复杂社会规范。为保障长对话训练稳定,引入分层优势估计机制,分别计算回合级与词元级优势。在SOTOPIA社交环境及狼人杀策略游戏中,训练模型展现出细腻的涌现式社交智能,如共情、说服与妥协寻求,证明即使在竞争场景下也能学习协作。尽管存在奖励滥用等实际挑战,结果表明丰富社交智能可在无监督条件下自然生成。本工作旨在激励未来对群体对话中人工智能社交智能的研究。
原文摘要 · Abstract (English)
This paper introduces OMAR: One Model, All Roles, a reinforcement learning framework that enables AI to develop social intelligence through multi-turn, multi-agent conversational self-play. Unlike traditional paradigms that rely on static, single-turn optimizations, OMAR allows a single model to role-play all participants in a conversation simultaneously, learning to achieve long-term goals and complex social norms directly from dynamic social interaction. To ensure training stability across long dialogues, we implement a hierarchical advantage estimation that calculates turn-level and token-level advantages. Evaluations in the SOTOPIA social environment and Werewolf strategy games show that our trained models develop fine-grained, emergent social intelligence, such as empathy, persuasion, and compromise seeking, demonstrating the effectiveness of learning collaboration even under competitive scenarios. While we identify practical challenges like reward hacking, our results show that rich social intelligence can emerge without human supervision. We hope this work incentivizes further research on AI social intelligence in group conversations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。