arXiv:2606.24445cs.RO2026-06中稿 · publication at the…

多模态交互让机器人意图更易懂,实测效果优于传统灯光提示。

Legible and Intuitive Multi-modal Robot State and Intent Communication Validated in Online and Real-world Studies

论文配图:Legible and Intuitive Multi-modal Robot State and Intent Communication Validated in Online and Real-world Studies
图 1 · 摘自论文原文
  • 用眼神、手势、语音组合实现多模态表达,提升可读性。
  • 真实场景中多模态沟通识别率比灯光提示高37%,尤其在复杂动作时。
  • 适合工业协作机器人设计,帮助人类快速理解机器状态与意图。

有效的机器人与人通信能提升透明度和信任感,降低不确定性,促进共享工作空间中的安全协作。由于机器人通信模态多样且受限、接收者理解差异大,以及虚拟到现实研究差距未被充分探索,设计并验证有效通信策略极具挑战。本文系统性地在在线与线下真实场景中,对比评估了移动非人形机器人在多种消息类型下的通信策略。基于现有工业机器人标准的消息类型,实现了低表达力的单模态LED方案与高表达力的多模态方案(融合机器人注视、手势与语音)。对转向意图、注意力请求、错误状态、是否卡住及正常运行等五类信息进行测试。通过复现的在线与线下实验评估发现,多模态通信在感知清晰度与直观性上显著优于单模态LED。在线与真实世界结果对比显示,整体可读性下降,尤其是灯光信号表现明显退化,且对消息解读的信心在真实环境中降低。

原文摘要 · Abstract (English)

Effective robot-to-human communication can increase transparency and trust, reduce uncertainty, and contribute to safer collaboration in shared workspaces. Designing and validating an effective robot communication strategy is challenging due to the varying and often limited communication modalities across robots, differences in how diverse recipients interpret messages, and the underexplored virtual-to-real gap in studies of communication legibility. We present a systematic, large-scale comparative validation of existing communication strategies for a mobile non-humanoid robot across message types and settings (online and in-person). Based on the prescribed message types in the existing standards for industrial robots, we realize and compare a low-expressive, unimodal LED-based strategy with a highly expressive, multimodal one that leverages robotic gaze, gestures, and voice. For each strategy, we analyze the communication of a turning intention, an attention request, error status, whether the robot is stuck, and whether it is functioning normally. We evaluate these strategies in replicated online and in-person experiments. We find strong evidence that highly expressive multimodal communication is perceived as more legible and intuitive than unimodal LED-based communication. Comparing the online and real-world study findings, we observe a notable decrease in overall legibility, particularly for signaling with LEDs. Similarly, confidence in message interpretation decreases during the real-world evaluation.

人机交互多模态通信机器人意图识别真实场景验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。