arXiv:2604.18982cs.AI2026-04ACL被引 2

用博弈论方法让对话智能体学会社交智慧,效果超越GPT-4o等大模型。

SAVOIR: Learning Social Savoir-Faire via Shapley-based Reward Attribution

  • 基于合作博弈论的沙普利值,公平分配每句话对对话结果的贡献。
  • 在SOTOPIA测试中7B模型性能媲美甚至超过GPT-4o和Claude-3.5-Sonnet。
  • 证明社交智能不同于逻辑推理,大模型未必擅长复杂人际互动。

社交智能——即在复杂人际互动中有效导航的能力——是语言代理面临的核心挑战。通过强化学习训练此类代理需解决信用分配问题:判断单个话语如何影响多轮对话的结果。现有方法直接使用语言模型分配回合级奖励,得出的归因具有回顾性且缺乏理论基础。我们提出SAVOIR(ShApley Value fOr SocIal RL),一种基于合作博弈论的新型原理性框架。该方法结合两种互补原则:期望效用变化将评估从回顾性归因转变为前瞻性估值,捕捉话语对未来有利轨迹的战略潜力;沙普利值则确保信用分配的公平性,并具备效率、对称性和边际性等公理化保障。在SOTOPIA基准上的实验表明,SAVOIR在所有评估设置下均达到新最佳性能,其7B模型表现与甚至超越多个专有模型,包括GPT-4o和Claude-3.5-Sonnet。值得注意的是,即使大型推理模型也持续表现不佳,表明社交智能所需能力与分析推理存在质的差异。

原文摘要 · Abstract (English)

Social intelligence, the ability to navigate complex interpersonal interactions, presents a fundamental challenge for language agents. Training such agents via reinforcement learning requires solving the credit assignment problem: determining how individual utterances contribute to multi-turn dialogue outcomes. Existing approaches directly employ language models to distribute episode-level rewards, yielding attributions that are retrospective and lack theoretical grounding. We propose SAVOIR (ShApley Value fOr SocIal RL), a novel principled framework grounded in cooperative game theory. Our approach combines two complementary principles: expected utility shifts evaluation from retrospective attribution to prospective valuation, capturing an utterance's strategic potential for enabling favorable future trajectories; Shapley values ensure fair credit distribution with axiomatic guarantees of efficiency, symmetry, and marginality. Experiments on the SOTOPIA benchmark demonstrate that SAVOIR achieves new state-of-the-art performance across all evaluation settings, with our 7B model matching or exceeding proprietary models including GPT-4o and Claude-3.5-Sonnet. Notably, even large reasoning models consistently underperform, suggesting social intelligence requires qualitatively different capabilities than analytical reasoning.

社交智能强化学习沙普利值对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。