arXiv:2604.03451cs.ROcs.CY2026-04

研究机器人如何用动作、灯光等信号让人类更好理解其行进意图。

Do Robots Need Body Language? Comparing Communication Modalities for Legible Motion Intent in Human-Shared Spaces

  • 测试动作、灯光、文字、语音四种信号对人类理解机器人意图的影响。
  • 发现显式信号(如文字)比隐式动作更准确,但动作能提升信任感。
  • 多模态一致信号效果更好,冲突信号会降低用户信心与信任。

在共享空间中,高自由度机器人常以难以解读的方式移动,迫使人类适应。这类机器人的运动常被感知为具有表现力,因此理解这些线索如何被人类解读至关重要。我们开展了一项在线视频研究,评估四种信号模态——表达性动作、灯光、文字和音频——对人类理解四足机器人(波士顿动力 Spot)即将进行的导航动作的影响。研究涵盖四个常见场景,测量各模态对人类(1)预测机器人下一步行动的准确性,(2)对该预测的信心,以及(3)对机器人安全行为的信任度的影响。研究还考察了表达性动作与显式信号的对比效果、多模态信号是否协同增强可理解性,以及矛盾信号如何影响用户信心与信任。本研究提供了关于隐式与显式信号策略相对有效性的初步证据。

原文摘要 · Abstract (English)

Robots in shared spaces often move in ways that are difficult for people to interpret, placing the burden on humans to adapt. High-DoF robots exhibit motion that people read as expressive, intentionally or not, making it important to understand how such cues are perceived. We present an online video study evaluating how different signaling modalities, expressive motion, lights, text, and audio, shape people's ability to understand a quadruped robot's upcoming navigation actions (Boston Dynamics Spot). Across four common scenarios, we measure how each modality influences humans' (1) accuracy in predicting the robot's next navigation action, (2) confidence in that prediction, and (3) trust in the robot to act safely. The study tests how expressive motions compare to explicit channels, whether aligned multimodal cues enhance interpretability, and how conflicting cues affect user confidence and trust. We contribute initial evidence on the relative effectiveness of implicit versus explicit signaling strategies.

人机交互机器人意图多模态通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。