机器人难发现对话错误,因用户常不主动反馈困惑。
Why Robots Are Bad at Detecting Their Mistakes: Limitations of Miscommunication Detection in Human-Robot Dialogue
- 用240组对话数据测试视觉模型,评估其识错能力
- 模型准确率仅略高于随机,远低于人类表现
- 用户即使觉察错误也常不表达,导致系统难以察觉
在人机交互中,检测沟通失误对维持用户参与度和信任至关重要。尽管人类能通过言语与非语言线索轻松识别交流错误,机器人却面临解读非语言反馈的挑战,即便计算机视觉在识别情绪表达方面已取得进展。本研究基于包含240组人机对话的多模态数据集,系统引入四类对话失败情境,评估前沿计算机视觉模型在误判检测中的表现。每轮对话后,用户反馈是否感知到错误,以分析模型准确性。结果显示,即使使用最先进的模型,其识别误判的能力仍仅略高于随机水平;而在更具情感表现力的数据集上,模型可成功识别出困惑状态。为探究原因,我们邀请人类评分者进行相同任务,他们同样只能识别约一半的诱导性误判,与模型表现一致。这一发现揭示了人机对话中一个根本性局限:即使用户感知到错误,也往往不会主动向机器人传达。该认知有助于合理预期视觉模型性能,并指导研究者设计更有效的互动机制,主动获取用户反馈。
原文摘要 · Abstract (English)
Detecting miscommunication in human-robot interaction is a critical function for maintaining user engagement and trust. While humans effortlessly detect communication errors in conversations through both verbal and non-verbal cues, robots face significant challenges in interpreting non-verbal feedback, despite advances in computer vision for recognizing affective expressions. This research evaluates the effectiveness of machine learning models in detecting miscommunications in robot dialogue. Using a multi-modal dataset of 240 human-robot conversations, where four distinct types of conversational failures were systematically introduced, we assess the performance of state-of-the-art computer vision models. After each conversational turn, users provided feedback on whether they perceived an error, enabling an analysis of the models' ability to accurately detect robot mistakes. Despite using state-of-the-art models, the performance barely exceeds random chance in identifying miscommunication, while on a dataset with more expressive emotional content, they successfully identified confused states. To explore the underlying cause, we asked human raters to do the same. They could also only identify around half of the induced miscommunications, similarly to our model. These results uncover a fundamental limitation in identifying robot miscommunications in dialogue: even when users perceive the induced miscommunication as such, they often do not communicate this to their robotic conversation partner. This knowledge can shape expectations of the performance of computer vision models and can help researchers to design better human-robot conversations by deliberately eliciting feedback where needed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。