让对话系统听懂用户反应,用语言反馈优化情感支持策略
Listening to the Echo: User-Reaction Aware Policy Optimization via Scalar-Verbal Hybrid Reinforcement Learning
- 用用户实时反应生成自然语言反馈,替代传统评分信号
- 在ESC和Sotopia数据集上显著提升正向情绪转化效果
- 适合需要精准情感响应的对话系统研究与开发
当前情感支持对话系统多依赖人工定义的标量奖励进行对齐,但此类信号信息稀疏,无法解释回复失败原因或适应动态用户状态,常偏离促进积极情绪转变的核心目标。实践中,用户在交互过程中的持续反应是最直接可靠的反馈来源。为此,我们提出反应感知策略优化框架RAPO,不基于评分规则而基于交互后果进行优化。RAPO将对话视为反应驱动过程,通过三个核心组件生成密集的自然语言反馈:事后对话选择,识别显著改变用户情绪轨迹的关键对话回合;生成式事后反馈,将用户反应转化为对比排序信号与自然语言批评;标量-语言混合策略优化,结合标量奖励优化全局对齐性与语言反馈蒸馏实现细粒度语义精炼。在ESC和Sotopia数据集上的大量实验表明,RAPO显著优于多个强基线强化学习方法,在推动正向交互结果方面表现更优。
原文摘要 · Abstract (English)
While current emotional support dialogue systems typically rely on expert-defined scalar rewards for alignment, these signals suffer from severe information sparsity. They cannot explain why a response failed or how to adapt to dynamic user states, often diverging from the actual goal of facilitating positive emotional shifts. In practice, the most direct and reliable learning signal emerges from the user's continuous reactions during ongoing interaction. We therefore propose Reaction Aware Policy Optimization (RAPO), a framework that optimizes over interaction consequences rather than rubric scores. RAPO treats dialogue as a reaction-driven process and utilizes simulated user responses to generate dense natural-language feedback through three core components: Hindsight Dialogue Selection, which isolates pivotal turns that meaningfully alter user emotional trajectories; Generative Hindsight Feedback, which transforms user reactions into contrastive ranking signals and natural-language critiques; and Scalar-Verbal Hybrid Policy Optimization, which couples scalar reward optimization for global alignment with verbal feedback distillation for fine-grained semantic refinement. Extensive experiments on ESC and Sotopia demonstrate that RAPO significantly outperforms strong reinforcement learning baselines in driving positive interaction outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。