arXiv:2505.24500cs.CLcs.AI2025-05

用时间感知的分层认知强化学习提升大模型社交智能

TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence

  • 引入时间感知分层认知框架,融合直觉与深度思考模式
  • 7B模型性能媲美DeepSeek-R1和OpenAI-O3,显著优于传统方法
  • 适用于提升大模型在社交场景下的推理与决策能力

大语言模型在数学、编程等需要缜密思维的任务中已取得显著进展,但在社交领域认知能力的提升,尤其是后训练阶段,仍研究不足。鉴于社交情境具有独特的时间线特征,且需融合直觉反应(系统1)与深层思考(系统2)的复合认知模式,而数学任务主要依赖系统2的逐步推理,本文提出时间感知的分层认知强化学习(TimeHC-RL)以增强大模型的社交智能。我们在八个具有不同数据模式的数据集上,通过五种后训练范式与两种测试时干预策略进行系统验证。实验表明,所提方法在性能上优于广泛采用的系统2强化学习方法,使7B规模模型实现显著跃升,其表现可比肩DeepSeek-R1与OpenAI-O3。此外,从后训练与测试时干预双视角的探索,揭示了多项关于提升大模型社交智能的关键洞见。

原文摘要 · Abstract (English)

Recently, Large Language Models (LLMs) have made significant progress in IQ-related domains that require careful thinking, such as mathematics and coding. However, enhancing LLMs' cognitive development in social domains, particularly from a post-training perspective, remains underexplored. Recognizing that the social world follows a distinct timeline and requires a richer blend of cognitive modes (from intuitive reactions (System 1) and surface-level thinking to deliberate thinking (System 2)) than mathematics, which primarily relies on System 2 cognition (careful, step-by-step reasoning), we introduce Temporal-aware Hierarchical Cognitive Reinforcement Learning (TimeHC-RL) for enhancing LLMs' social intelligence. In our experiments, we systematically explore improving LLMs' social intelligence and validate the effectiveness of the TimeHC-RL method, through five other post-training paradigms and two test-time intervention paradigms on eight datasets with diverse data patterns. Experimental results reveal the superiority of our proposed TimeHC-RL method compared to the widely adopted System 2 RL method. It gives the 7B backbone model wings, enabling it to rival the performance of advanced models like DeepSeek-R1 and OpenAI-O3. Additionally, the systematic exploration from post-training and test-time interventions perspectives to improve LLMs' social intelligence has uncovered several valuable insights.

大模型社交智能强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。