arXiv:2508.12935cs.AI2025-08被引 5

用未来导向奖励训练大模型,让聊天机器人更懂如何长期安慰人。

Towards Open-Ended Emotional Support Conversations in LLMs via Reinforcement Learning with Future-Oriented Rewards

  • 通过模拟未来对话轨迹获取情感支持奖励,端到端训练响应策略。
  • 在两个公开数据集上,系统在完成目标和回复质量上均超越现有方法。
  • 适合研究情感对话、强化学习应用或想提升聊天机器人共情能力的人。

情感支持对话(ESC)系统旨在缓解用户情绪困扰,并提供长期、系统的心理支持。然而,大多数基于大语言模型(LLM)的ESC系统依赖预设策略,在复杂真实场景中效果受限。为实现对多样化情绪问题的灵活应对,本文提出一种全新的端到端框架RLFF-ESC,直接通过强化学习习得持久的情感支持响应能力。为实现持续支持,我们首先采用基于LLM的多智能体机制模拟未来对话轨迹并收集未来导向奖励;随后训练一个未来导向奖励模型,用于指导情感支持策略模型的优化。此外,我们在生成回复时引入显式推理过程,进一步提升回应的质量、相关性与情境适配性。我们在Qwen2.5-7B-Instruct-1M和LLaMA3.1-8B-Instruct模型上评估主干策略模型,测试所提框架在两个公开ESC数据集上的表现。实验结果表明,RLFF-ESC在目标达成率与回复质量上均持续优于现有基线。

原文摘要 · Abstract (English)

Emotional Support Conversation (ESC) systems aim to alleviate users' emotional difficulties and provide long-term, systematic support for emotional well-being. However, most large language model (LLM)-based ESC systems rely on predefined strategies, which limits their effectiveness in complex, real-life scenarios. To enable flexible responses to diverse emotional problem scenarios, this paper introduces a novel end-to-end framework (RLFF-ESC) that directly learns enduring emotionally supportive response skills using reinforcement learning. For sustained emotional support, we first employ an LLM-based multi-agent mechanism to simulate future dialogue trajectories and collect future-oriented rewards. We then train a future-oriented reward model, which is subsequently used to train the emotional support policy model. Additionally, we incorporate an explicit reasoning process during response generation to further enhance the quality, relevance, and contextual appropriateness of the system's responses. We evaluate the backbone policy model on Qwen2.5-7B-Instruct-1M and LLaMA3.1-8B-Instruct models, testing the proposed RLFF-ESC framework across two public ESC datasets. Experimental results demonstrate that RLFF-ESC consistently outperforms existing baselines in terms of goal completion and response quality.

情感对话强化学习大模型共情生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。