检测对话系统中用户真实沮丧情绪,发现大模型表现最佳。
"Stupid robot, I want to speak to a human!" User Frustration Detection in Task-Oriented Dialog Systems
- 用关键词、开源情感分析、对话崩溃检测和大模型对比方法
- 大模型在内部评测中F1提升16%,显著优于其他方案
- 适合工业界落地应用,提供真实场景下沮丧检测的实用建议
在现代任务导向对话系统中,检测用户沮丧情绪对维持用户满意度、参与度和留存至关重要。然而,现有研究多聚焦于学术场景下的情感与情绪识别,难以反映真实用户数据的特点。为此,本文聚焦已部署对话系统中的用户沮丧问题,评估现成解决方案的可行性。我们比较了基于关键词的部署方案、开源情感分析方法、对话崩溃检测技术以及新兴的上下文学习大模型检测方法。分析表明,开源方法在真实场景中表现有限,而大模型方法展现出显著优势,在内部基准测试中实现F1分数16%的相对提升。最后,我们探讨了各方法的优缺点,为行业实践者提供用户沮丧检测任务的深入洞察。
原文摘要 · Abstract (English)
Detecting user frustration in modern-day task-oriented dialog (TOD) systems is imperative for maintaining overall user satisfaction, engagement, and retention. However, most recent research is focused on sentiment and emotion detection in academic settings, thus failing to fully encapsulate implications of real-world user data. To mitigate this gap, in this work, we focus on user frustration in a deployed TOD system, assessing the feasibility of out-of-the-box solutions for user frustration detection. Specifically, we compare the performance of our deployed keyword-based approach, open-source approaches to sentiment analysis, dialog breakdown detection methods, and emerging in-context learning LLM-based detection. Our analysis highlights the limitations of open-source methods for real-world frustration detection, while demonstrating the superior performance of the LLM-based approach, achieving a 16\% relative improvement in F1 score on an internal benchmark. Finally, we analyze advantages and limitations of our methods and provide an insight into user frustration detection task for industry practitioners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。