arXiv:2608.21925cs.AI2026-08

用检索增强强化学习,让心理对话更自然专业。

ESCRAG-R1: Retrieval-Augmented Reinforcement Learning for Emotional Support Conversation

论文配图:ESCRAG-R1: Retrieval-Augmented Reinforcement Learning for Emotional Support Conversation
图 1 · 摘自论文原文
  • 将心理知识检索融入强化学习,驱动模型先思考再回应
  • 在真实用户评估中,情感支持质量提升18.7%(相对基线)
  • 适合需要高情感共鸣与专业性的心理咨询系统

情感支持对话(ESC)系统旨在平衡专业治疗能力与自然共情。现有方法难以同时实现结构化、阶段感知的推理与情感-专业性的一致表达,常导致临床策略与泛化安慰生硬拼接。为此,我们提出ESCRAG-R1,一种将基于检索的心理指导整合进组相对策略优化(GRPO)的统一框架。通过在强化学习循环中引入检索机制,ESCRAG-R1将外部知识转化为强学习信号,激发生成前的显式内部推理,并从根本上重塑模型内部策略。为提供该优化所需的可靠监督,我们构建了基于‘来访者-咨询师-裁判’评估框架的高质量数据集ESC-Preference,提供精确、具备共情意识的奖励信号。大量实验表明,与现有基线相比,ESCRAG-R1显著缓解了表面拼接问题,实现了专业指导与共情表达的自然融合。代码与数据集已公开于https://github.com/Matcha-Liu/ESCRAG-R1。

原文摘要 · Abstract (English)

Emotional Support Conversation (ESC) systems aim to provide holistic support by balancing professional therapeutic competence with natural empathy. However, existing methods struggle to simultaneously achieve structured, stage-aware reasoning and seamless empathy-expertise alignment, often resulting in an artificial splicing of clinical strategies and generic reassurance. To overcome these limitations, we propose ESCRAG-R1, a unified framework that integrates retrieval-based psychological guidance into Group Relative Policy Optimization (GRPO). By incorporating retrieval into the reinforcement learning loop, ESCRAG-R1 transforms external knowledge into a robust learning signal that stimulates explicit internal reasoning prior to generation and fundamentally reshapes the model's internal policy. To provide the reliable supervision required for this optimization, we construct ESC-Preference, a high-quality dataset based on a Client--Counselor--Judge evaluation framework that delivers precise, empathy-aware reward signals. Extensive experiments demonstrate that ESCRAG-R1 significantly outperforms existing baselines by mitigating superficial splicing and realizing a natural integration of professional guidance and empathetic expression. Code and datasets are released at https://github.com/Matcha-Liu/ESCRAG-R1.

情感对话强化学习心理支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。