通过可验证情绪反馈实现对话策略自进化,提升多轮共情对话能力。
Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

- 双环架构:内环优化对话策略,外环动态调整训练体验分布。
- 在SAGE数据集上,Qwen3-8B整体得分从53.87提升至79.24。
- 适用于需要长期共情支持的对话系统研发与评估。
大语言模型虽具对话能力,但共情表现仍具挑战。共情支持本质上是多轮且路径依赖的:用户逐步披露关切,情绪随时间演变,早期回应影响信任与接受度。基于可验证情绪奖励的强化学习为长时交互提供可扩展监督。然而,现有方法在固定训练交互分布下演化对话策略,导致策略能力与训练经验不匹配。本文提出一种由可验证情绪反馈驱动的双环自进化框架。在用户模拟器和验证器冻结的前提下,内环利用连续情绪奖励优化多轮策略,外环则利用相同结果估计策略相对交互效用并适应训练体验。为从稀疏、随机的滚动中获得估计,框架在每组内保持情景与交互状态不变,并优先选择组通过率接近策略能力边界的条件。层级控制器在不同支持意图间共享证据,不确定性引导探索与均匀重放防止过早排除。最终生成的分布产生轨迹,闭环运行且不增加采样预算。在SAGE数据集上,该框架将Qwen3-8B整体得分从53.87提升至79.24,较协议匹配的均匀情绪奖励强化学习高出7.23分。
原文摘要 · Abstract (English)
Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently multi-turn and path-dependent: users disclose concerns gradually, emotions evolve over time, and early responses shape trust and receptivity. Reinforcement learning with verifiable emotion rewards provides scalable supervision for long-horizon interactions. However, existing methods evolve the dialogue policy while keeping its training interaction distribution fixed, creating a mismatch between policy competence and training experience. We introduce a dual-loop self-evolution framework driven by verifiable emotion feedback. With the user simulator and verifier frozen, the inner loop optimizes the multi-turn policy using continuous emotion rewards, while the outer loop uses the same outcomes to estimate policy-relative interaction utility and adapt experience. To obtain estimates from sparse, stochastic rollouts, the framework holds the scenario and interaction state constant within each group and prioritizes conditions whose group pass rates lie near the policy's competence boundary. A hierarchical controller shares evidence across support intents, while uncertainty-guided exploration and uniform rehearsal prevent premature exclusion. The resulting distribution generates trajectories, closing both loops without increasing the rollout budget. On SAGE, our framework raises Qwen3-8B Overall from 53.87 to 79.24 and outperforms protocol-matched uniform emotion-reward reinforcement learning by 7.23 points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。