用视觉语言模型让机器人实时解释导航决策,提升人机信任。
Trust Through Transparency: Explainable Social Navigation for Autonomous Mobile Robots via Vision-Language Models
- 结合视觉语言模型与热力图生成自然语言解释
- 30人实验显示多数用户偏好实时说明,信任度上升
- 通过混淆矩阵验证解释与人类预期的一致性
服务型与辅助型机器人正越来越多地部署于动态社交环境,但确保交互过程的透明性与可解释性仍是重大挑战。本文提出一种多模态可解释性模块,融合视觉语言模型与热力图,实现机器人在导航过程中对环境感知、分析与决策的自然语言表述。用户研究(n=30)表明,多数参与者更偏好实时解释,反映出信任与理解的提升。通过混淆矩阵分析验证了系统输出与人类预期之间的匹配程度。实验与仿真结果均证明,可解释性显著提升了自主导航中的信任度与可理解性。
原文摘要 · Abstract (English)
Service and assistive robots are increasingly being deployed in dynamic social environments; however, ensuring transparent and explainable interactions remains a significant challenge. This paper presents a multimodal explainability module that integrates vision language models and heat maps to improve transparency during navigation. The proposed system enables robots to perceive, analyze, and articulate their observations through natural language summaries. User studies (n=30) showed a preference of majority for real-time explanations, indicating improved trust and understanding. Our experiments were validated through confusion matrix analysis to assess the level of agreement with human expectations. Our experimental and simulation results emphasize the effectiveness of explainability in autonomous navigation, enhancing trust and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。