让机器人理解语音与视觉的关联,提升人机交互的可解释性。
Enhancing Explainability with Multimodal Context Representations for Smarter Robots
- 构建多模态联合表征与时间对齐模块,融合语音与视觉信息。
- 实现用户话语与场景感知的实时相关性评估,支持动态推理。
- 适用于需高透明度的人机协作场景,如医疗、教育机器人。
近年来人工智能显著进步,推动了机器人领域的创新。尽管机器人能以日益自主的方式执行复杂任务,但在确保可解释性和以用户为中心的设计方面仍面临挑战。人机交互(HRI)中的关键问题在于让机器人能够有效感知和推理多模态输入(如语音与视觉),以建立信任并实现无缝协作。本文提出一种通用且可解释的多模态上下文表示框架,旨在改进语音与视觉模态的融合。我们以评估用户言语与机器人视觉感知之间的“相关性”为例,提出包含多模态联合表征模块与时间对齐模块的方法,使机器人能够通过时序对齐来判断输入间的相关性。最后,讨论该框架在提升人机交互可解释性方面的多方面潜力。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) has significantly advanced in recent years, driving innovation across various fields, especially in robotics. Even though robots can perform complex tasks with increasing autonomy, challenges remain in ensuring explainability and user-centered design for effective interaction. A key issue in Human-Robot Interaction (HRI) is enabling robots to effectively perceive and reason over multimodal inputs, such as audio and vision, to foster trust and seamless collaboration. In this paper, we propose a generalized and explainable multimodal framework for context representation, designed to improve the fusion of speech and vision modalities. We introduce a use case on assessing 'Relevance' between verbal utterances from the user and visual scene perception of the robot. We present our methodology with a Multimodal Joint Representation module and a Temporal Alignment module, which can allow robots to evaluate relevance by temporally aligning multimodal inputs. Finally, we discuss how the proposed framework for context representation can help with various aspects of explainability in HRI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。