通过预测用户满意度提升对话系统主动性,解决罕见语句和冷门领域识别难题。
Reward-Driven Interaction: Enhancing Proactive Dialogue Agents through User Satisfaction Prediction
- 引入对比自监督与领域意图分类两个辅助任务,增强用户语句与会话表征学习。
- 在DuerOS数据集上,罕见语句错误识别准确率显著提升,长尾领域表现改善明显。
- 适合需要高鲁棒性主动对话系统的工业场景,尤其关注低频查询与语音识别错误。
奖励驱动的主动对话系统需精准估计用户满意度作为内在奖励信号,以制定最优交互策略。该框架在工业对话系统中检测到潜在用户不满时触发澄清问题。传统方法依赖基于用户行为后验标签训练神经网络,但存在两大局限:(1) 噪声奖励监督,依赖事后用户动作生成的弱标签会引入偏差,难以捕捉因语音识别错误引发的满意度信号;(2) 长尾反馈稀疏,用户查询遵循幂律分布,导致低频领域预测精度下降。噪声标签与长尾分布使模型难以学习有效的用户语句与会话表征。为此,本文提出两个辅助任务:一是对比自监督学习任务,帮助模型学习罕见语句表征并识别语音识别错误;二是领域-意图分类任务,增强模型对长尾领域会话的理解能力。所提方法在DuerOS数据集上验证,显著提升了罕见语句错误识别准确率及长尾领域表现。
原文摘要 · Abstract (English)
Reward-driven proactive dialogue agents require precise estimation of user satisfaction as an intrinsic reward signal to determine optimal interaction strategies. Specifically, this framework triggers clarification questions when detecting potential user dissatisfaction during interactions in the industrial dialogue system. Traditional works typically rely on training a neural network model based on weak labels which are generated by a simple model trained on user actions after current turn. However, existing methods suffer from two critical limitations in real-world scenarios: (1) Noisy Reward Supervision, dependence on weak labels derived from post-hoc user actions introduces bias, particularly failing to capture satisfaction signals in ASR-error-induced utterances; (2) Long-Tail Feedback Sparsity, the power-law distribution of user queries causes reward prediction accuracy to drop in low-frequency domains. The noise in the weak labels and a power-law distribution of user utterances results in that the model is hard to learn good representation of user utterances and sessions. To address these limitations, we propose two auxiliary tasks to improve the representation learning of user utterances and sessions that enhance user satisfaction prediction. The first one is a contrastive self-supervised learning task, which helps the model learn the representation of rare user utterances and identify ASR errors. The second one is a domain-intent classification task, which aids the model in learning the representation of user sessions from long-tailed domains and improving the model's performance on such domains. The proposed method is evaluated on DuerOS, demonstrating significant improvements in the accuracy of error recognition on rare user utterances and long-tailed domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。