分离策略与回应,让大模型更懂心理支持。
DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimization
- 将情绪支持拆分为策略规划和共情回应两步训练
- 在偏好数据集上实现更低的偏见和更高的回应质量
- 适合需要精准心理支持的AI对话系统开发者
近期情感支持对话(ESC)研究通过监督微调(SFT)改进了大语言模型的情感支持生成能力,但常见心理错误仍存在。尽管直接偏好优化(DPO)在成对偏好学习中展现出减少错误的潜力,但在ESC任务中的效果受限于两个关键挑战:(1)数据结构纠缠:现有ESC数据天然混合心理策略与回应内容,难以构建高质量偏好对;(2)优化模糊性:直接应用原始DPO处理此类纠缠数据导致训练目标模糊。为此,我们提出推理偏好挖掘(IPM),构建高质量偏好数据,形成IPM-PrefDial数据集。基于此,我们设计解耦式ESC框架,受格罗斯情绪调节扩展模型启发,将任务分解为策略规划与共情回应两个连续子任务,分别经由SFT训练后,再通过DPO对齐心理偏好。大量实验表明,所提框架优于联合优化基线,显著降低偏好偏差并提升回应质量。
原文摘要 · Abstract (English)
Recent advances in Emotional Support Conversation (ESC) have improved emotional support generation by fine-tuning Large Language Models (LLMs) via Supervised Fine-Tuning (SFT). However, common psychological errors still persist. While Direct Preference Optimization (DPO) shows promise in reducing such errors through pairwise preference learning, its effectiveness in ESC tasks is limited by two key challenges: (1) Entangled data structure: Existing ESC data inherently entangles psychological strategies and response content, making it difficult to construct high-quality preference pairs; and (2) Optimization ambiguity: Applying vanilla DPO to such entangled pairwise data leads to ambiguous training objectives. To address these issues, we introduce Inferential Preference Mining (IPM) to construct high-quality preference data, forming the IPM-PrefDial dataset. Building upon this data, we propose a Decoupled ESC framework inspired by Gross's Extended Process Model of Emotion Regulation, which decomposes the ESC task into two sequential subtasks: strategy planning and empathic response generation. Each was trained via SFT and subsequently enhanced by DPO to align with the psychological preference. Extensive experiments demonstrate that our Decoupled ESC framework outperforms joint optimization baselines, reducing preference bias and improving response quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。