为大模型社交智能设计分句多维奖励,提升谈判协作效果
Sotopia-RL: Reward Design for Social Intelligence
- 将任务结果分解到每句话,精准分配奖励
- 在Sotopia-hard上达成7.17分,超越现有方法
- 适合研究社交推理与强化学习的开发者
社交智能已成为大型语言模型的关键能力,使其能有效参与协作与谈判等现实社会任务。强化学习天然适合训练社交智能体,因其可让模型通过社会互动直接学习复杂策略,无需人工标注。然而,社交任务存在两大特性:(1)单句质量与最终成功无严格关联;(2)成功需多维度评估标准。因此,我们主张设计分句级、多维度奖励模型以支持强化学习训练。为此,提出Sotopia-RL框架,将粗粒度的回合级反馈细化为分句级、多维度奖励。分句级信用分配将结果归因于具体语句,多维奖励捕捉社交互动的丰富性并减少奖励黑客问题。在开放式的社交学习环境Sotopia中的实验表明,Sotopia-RL在Sotopia-hard上达到7.17分,Sotopia-full上达8.31分,显著优于现有方法。消融实验证明,分句级信用分配与多维奖励设计对强化学习训练均不可或缺。
原文摘要 · Abstract (English)
Social intelligence has become a critical capability for large language models (LLMs), enabling them to engage effectively in real-world social tasks such as collaboration and negotiation. Reinforcement learning (RL) is a natural fit for training socially intelligent agents because it allows models to learn sophisticated strategies directly through social interactions without requiring human annotations. However, there are two unique parts about social intelligence tasks: (1) the quality of individual utterances in social interactions is not strictly related to final success; (2) social interactions require multi-dimensional rubrics for success. Therefore, we argue that it is necessary to design rewards for building utterance-level multi-dimensional reward models to facilitate RL training for social intelligence tasks. To address these challenges, we propose Sotopia-RL, a novel framework that refines coarse episode-level feedback into utterance-level, multi-dimensional rewards. Utterance-level credit assignment attributes outcomes to individual utterances, while multi-dimensional rewards capture the full richness of social interactions and reduce reward hacking. Experiments in Sotopia, an open-ended social learning environment, demonstrate that Sotopia-RL achieves state-of-the-art social goal completion scores (7.17 on Sotopia-hard and 8.31 on Sotopia-full), significantly outperforming existing approaches. Ablation studies confirm the necessity of both utterance-level credit assignment and multi-dimensional reward design for RL training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。