用评分标准引导强化学习,让大模型更懂人的情感与个性。
Kardia-R1: Unleashing LLMs to Reason toward Understanding and Empathy for Emotional Support via Rubric-as-Judge Reinforcement Learning
- 用可解释的评分规则指导模型逐步推理情感与身份
- 在17.8万组对话中验证,模型在共情和一致性上显著提升
- 适合研究情感计算、个性化对话系统的学者与开发者
随着网络平台向个性化和情感复杂化演进,对话代理需超越表面共情,实现基于身份认知的情绪推理。现有系统存在两大缺陷:(1)依赖以情境为中心的数据集,缺乏持续用户身份信息,难以捕捉个性化情感细节;(2)依赖模糊、粗糙的奖励信号,阻碍可验证的共情推理发展。为此,我们提出KardiaBench,一个大规模用户基础基准,包含178,080个问答对,覆盖22,080轮多轮对话,锚定671个真实用户画像。该数据集通过模型在环的迭代式评分引导优化流程构建,确保心理合理性与人格一致性。在此基础上,我们提出Kardia-R1框架,采用基于GRPO的“评分标准即裁判”共情强化学习(Rubric-ERL),通过可解释的人类对齐评分奖励,紧密耦合用户理解、情绪推断与支持性回复生成。在四个LLM基线上的实验表明,Kardia-R1在情绪准确率、共情度、相关性、人格一致性与安全性上均优于其他方法。数据集与模型将开源于https://github.com/JhCircle/Kardia-R1。
原文摘要 · Abstract (English)
As web platforms evolve towards greater personalization and emotional complexity, conversational agents must transcend superficial empathy to demonstrate identity-aware emotional reasoning. However, existing systems face two limitations: (1) reliance on situation-centric datasets lacking persistent user identity, which hampers the capture of personalized affective nuances; and (2) dependence on opaque, coarse reward signals that hinder development of verifiable empathetic reasoning. To address these gaps, we introduce KardiaBench, a large-scale user-grounded benchmark comprising 178,080 QA pairs across 22,080 multi-turn conversations anchored to 671 real-world profiles. The dataset is constructed via a model-in-the-loop pipeline with iterative rubric-guided refinement to ensure psychological plausibility and persona consistency. This progressive empathy pipeline that integrates user comprehension, contextual reasoning, and emotion perception into conversations, followed by iterative critique and rubric-based refinement to ensure psychological plausibility, emotional fidelity, and persona consistency. Building on this, we propose Kardia-R1, a framework that trains models for interpretable, stepwise empathetic cognition. Kardia-R1 leverages Rubric-as-Judge Empathetic Reinforcement Learning (Rubric-ERL), a GRPO-based method that uses explainable, human-aligned rubric rewards to tightly couple user understanding, emotional inference, and supportive response generation. Extensive experiments across four LLM backbones demonstrate that Kardia-R1 consistently outperforms othet methods in emotion accuracy, empathy, relevance, persona consistency, and safety. Our dataset and model will be released at https://github.com/JhCircle/Kardia-R1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。