用分层推理和动态奖励提升大模型社交智能
Enhancing Social Intelligence in LLMs with Hierarchical Reasoning and Utterance-Level Goal Rewarding

- 分两阶段处理对话:先定策略,再生成语言
- 在SOTOPA上使目标达成率超GPT-4o 7.32%
- 适合研究多智能体社交协作的开发者
大型语言模型在结构化任务中表现优异,但在动态社交互动中表现不佳,因缺乏长期目标协调与快速适应能力。现有方法对每句话使用统一目标奖励,忽略每轮对话的具体目标差异及策略逻辑。受计划行为理论启发,我们提出Think-Strategy-Response(TSR)框架,将社交对话分解为高层战略规划与底层语言执行两个层次。为优化TSR,引入线性化分层强化学习与方差门控奖励(LHRL-VGR)算法,根据目标达成分数方差动态分配奖励,平衡目标完成与策略一致性。在SOTOPA基准测试中,该方法微调Qwen2.5-7B模型,在多智能体社交协商任务中实现7.32%的胜率超越GPT-4o,达到当前最优水平。
原文摘要 · Abstract (English)
Large language models (LLMs) excel in structured tasks but struggle with dynamic social interactions, where success requires long-term goal coordination and rapid adaptation. Current methods often apply uniform goal-based rewards to every utterance, overlooking the specificity of objectives at each dialogue turn and failing to account for the rationale of potential strategies. Inspired by the Theory of Planned Behavior, we propose the Think-Strategy-Response (TSR) framework, which decomposes social dialogue into two hierarchical stages: high-level strategic planning and low-level linguistic execution. To optimize TSR, we introduce Linearized Hierarchical Reinforcement Learning with Variance-Gated Rewards (LHRL-VGR), a novel algorithm that dynamically routes rewards - balancing goal completion and strategy adherence - based on the variance of goal achievement scores. Experiments on the SOTOPIA benchmark show that our approach fine-tunes a Qwen2.5-7B agent to surpass the GPT-4o baseline by 7.32% in goal completion success, demonstrating state-of-the-art performance in multi-agent social negotiation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。