让语言智能体学会根据社交场景自动调节思考深度,更高效地协作谈判。
Adaptive Social Learning via Mode Policy Optimization for Language Agents
- 设计多粒度思考模式,从直觉反应到深度推理分层响应。
- 在社交任务中实现15.6%性能提升,思考链缩短32.8%。
- 适合需要动态决策的对话、谈判类智能体研究者参考。
有效的社交智能模拟要求语言智能体在动态交互中动态调整推理深度,而现有方法或缺乏显式推理,或对所有场景统一使用冗长的思维链,导致令牌消耗过高且行为僵化。本文提出自适应社交学习(ASL)框架,旨在提升语言智能体在复杂社交互动中的自适应推理能力。基于认知控制理论,我们识别出从直觉回应到深度思辨的多层次推理模式,并设计了自适应模式策略优化(AMPO)算法,实现上下文感知的模式切换与推理。ASL在基准社交智能环境上表现优异:相较GPT-4o提升15.6%任务性能;相较于GRPO,AMPO性能高7.0%,且思维链长度减少32.8%,验证了其高效性与自适应优势。
原文摘要 · Abstract (English)
Effective social intelligence simulation requires language agents to dynamically adjust reasoning depth, a capability notably absent in current studies. Existing methods either lack explicit reasoning or employ lengthy Chain-of-Thought reasoning uniformly across all scenarios, resulting in excessive token usage and inflexible social behaviors in tasks such as negotiation or collaboration. To address this, we propose an $\textbf{A}$daptive $\textbf{S}$ocial $\textbf{L}$earning ($\textbf{ASL}$) framework in this paper, aiming to improve the adaptive reasoning ability of language agents in dynamic social interactions. To this end, we first identify the hierarchical reasoning modes under such context, ranging from intuitive response to deep deliberation based on the cognitive control theory. We then develop the $\textbf{A}$daptive $\textbf{M}$ode $\textbf{P}$olicy $\textbf{O}$ptimization ($\textbf{AMPO}$) algorithm to learn the context-aware mode adaptation and reasoning. Our framework advances existing research in three key aspects: (1) Multi-granular reasoning mode design, (2) Context-aware mode switching in rich social interaction, and (3) Token-efficient reasoning with depth adaptation. Extensive experiments on the benchmark social intelligence environment verify that ASL achieves 15.6% higher task performance than GPT-4o. Notably, our AMPO outperforms GRPO by 7.0% with 32.8% shorter thinking chains, demonstrating the advantages of our AMPO and the learned adaptive reasoning ability over GRPO's solution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。