用世界模型预测用户情绪意图,让对话系统更懂人。
Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling
- 构建对话世界模型,预测用户情绪、情感和意图
- 在多个数据集上实现情绪识别与情感分析的领先表现
- 适合需要共情能力的对话场景,如心理支持类应用
世界模型已在机器人、游戏和自动驾驶中广泛应用,但在自然语言任务中仍较有限。本文构建了对话世界模型,可预测用户的语气、情感和意图及未来话语。通过定义部分可观测马尔可夫决策过程(POMDP),将情绪、情感和意图建模为用户信念,并通过最大化信息瓶颈来求解。基于此信念建模,我们提出了一个基于模型的强化学习框架DreamCUB。实验表明,预训练的对话世界模型在情绪分类和情感识别任务上达到当前最优性能;联合训练策略、价值函数与世界模型后,对话质量显著提升。进一步分析显示,该方法具备合理的探索-利用平衡,且在跨领域场景(如共情对话)中具有良好迁移能力。
原文摘要 · Abstract (English)
World models have been widely utilized in robotics, gaming, and auto-driving. However, their applications on natural language tasks are relatively limited. In this paper, we construct the dialogue world model, which could predict the user's emotion, sentiment, and intention, and future utterances. By defining a POMDP, we argue emotion, sentiment and intention can be modeled as the user belief and solved by maximizing the information bottleneck. By this user belief modeling, we apply the model-based reinforcement learning framework to the dialogue system, and propose a framework called DreamCUB. Experiments show that the pretrained dialogue world model can achieve state-of-the-art performances on emotion classification and sentiment identification, while dialogue quality is also enhanced by joint training of the policy, critic and dialogue world model. Further analysis shows that this manner holds a reasonable exploration-exploitation balance and also transfers well to out-of-domain scenarios such as empathetic dialogues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。