基于心理学的双向共情评分模型,提升大模型的情感支持能力
PERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models
- 从支持者与求助者双向视角分解共情,引入旁观者监控互动质量
- 在基准测试和工业数据集上均超越现有方法超10%,用户偏好达70%
- 适合需要高情感智能的对话系统研发,如心理陪伴、客服场景
大型语言模型在以人为本的应用中日益普及,但常无法提供实质性的心理支持。尽管强化学习已被用于提升模型共情能力,但现有奖励模型通常仅从单一角度评估共情,忽略了共情循环理论所定义的支持者与求助者之间本有的双向互动特性。为此,我们提出心理学基础的共情奖励建模(PERM)。PERM通过双向分解实现共情评估:1)支持者视角,衡量内在共鸣与表达传递;2)求助者视角,评估情绪接收效果;此外还引入旁观者视角以监控整体互动质量。在广泛使用的情绪智力基准和工业级日常对话数据集上的大量实验表明,PERM性能优于当前最优基线超过10%。此外,盲测用户研究显示70%用户更偏好该方法,凸显其生成更具共情响应的有效性。代码、数据集与模型已开源于https://github.com/ZhengWwwq/PERM。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in human-centric applications, yet they often fail to provide substantive emotional support. While Reinforcement Learning (RL) has been utilized to enhance empathy of LLMs, existing reward models typically evaluate empathy from a single perspective, overlooking the inherently bidirectional interaction nature of empathy between the supporter and seeker as defined by Empathy Cycle theory. To address this limitation, we propose Psychology-grounded Empathetic Reward Modeling (PERM). PERM operationalizes empathy evaluation through a bidirectional decomposition: 1) Supporter perspective, assessing internal resonation and communicative expression; 2) Seeker perspective, evaluating emotional reception. Additionally, it incorporates a bystander perspective to monitor overall interaction quality. Extensive experiments on a widely-used emotional intelligence benchmark and an industrial daily conversation dataset demonstrate that PERM outperforms state-of-the-art baselines by over 10\%. Furthermore, a blinded user study reveals a 70\% preference for our approach, highlighting its efficacy in generating more empathetic responses. Our code, dataset, and models are available at https://github.com/ZhengWwwq/PERM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。