arXiv:2508.09521cs.CLcs.AI2025-08被引 1

让对话助手更懂共情:分三步推理再回应,效果更好更自然。

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning

  • 将共情对话拆解为分析历史、识情绪、选策略三步,结构化推理
  • 在多轮对话中提升共情度17.3%、策略契合度23.1%,且减少重复回复
  • 适合需要高情感智能的客服、心理咨询等场景

情感支持对话不仅需要流畅回复,还需理解对方处境与情绪,选择合适策略,并以自然、类人方式回应。尽管大语言模型已有进展,现有系统仍缺乏心理学指导的结构化推理能力。同时,强化学习因奖励信号不可靠而难以有效优化,且微调易放大重复性回复。为此,我们提出结构化共情推理框架,将支持过程分为三步:对话历史分析、多模态情绪状态推断、策略选择,再生成最终回复。为此构建了SER数据集,包含逐步骤正确性标注与成对回复偏好。进一步提出PEER,采用GRPO与统一过程-结果奖励模型(UnifiReward),联合评估推理步骤与最终回复质量。为降低重复,引入基于个性的重写增强数据,并对冗余输出进行降权。大量实验表明,该方法在不牺牲多样性前提下,显著提升共情度、策略一致性与人类自然度。

原文摘要 · Abstract (English)

Emotional support conversations require more than fluent responses. Supporters need to understand the seeker's situation and emotions, adopt an appropriate strategy, and respond in a natural, human-like manner. Despite advances in large language models, current systems often lack structured, psychology-informed reasoning. Additionally, it is challenging to enhance these systems through reinforcement learning because of unreliable reward signals. Moreover, reinforcement fine-tuning can amplify repetitive response patterns. We propose structured empathetic reasoning, which breaks support into three steps: conversation history analysis, multimodal emotional state inference, and strategy selection, prior to generating the final reply. To implement this, we introduce SER, a fine-grained dataset with step-level correctness labels and pairwise response preferences. We then present PEER, which uses GRPO with UnifiReward, a unified process-outcome reward model for evaluating both reasoning steps and final responses in multi-turn interactions. To reduce repetition, we enhance data with personality-based rewriting and down-weight redundant outputs. Comprehensive experiments show improved empathy, strategy alignment, and human-likeness without sacrificing diversity. Code and data are available at https://github.com/Yunxiao-Wang/PEER.

共情对话强化学习结构化推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。