arXiv:2601.14587cs.HCcs.RO2026-01被引 2

用AR界面实时展示机器人能做什么、不能做什么,提升人机协作透明度。

Explainable OOHRI: Communicating Robot Capabilities and Limitations as Augmented Reality Affordances

  • 通过视觉符号和动态菜单在AR中直观呈现机器人的动作能力与限制。
  • 用户能准确理解机器人局限,并在任务中主动调整指令。
  • 适合需要人机协作的复杂场景,如工业装配或家庭服务机器人。

人类交互对发出个性化指令和在机器人可能出错时提供协助至关重要。然而,机器人仍像黑箱,难以让用户了解其动态的能力与限制。为此,我们提出可解释的对象导向人机交互(X-OOHRI),一种基于增强现实(AR)的界面,通过视觉标记、径向菜单、颜色编码和解释标签传达机器人动作的可能性与约束。系统利用视觉-语言模型将物体属性与机器人限制编码为面向对象的结构,实现在模拟环境中对虚拟孪生体进行空间对齐的实时解释生成与直接操作。我们已将端到端流程集成至实体机器人,并展示了从低层级抓取放置到高层级指令的多样化应用场景。最后,通过用户研究发现,参与者能够有效发出面向对象的命令,形成对机器人局限的准确心理模型,并实现混合主动性协商。

原文摘要 · Abstract (English)

Human interaction is essential for issuing personalized instructions and assisting robots when failure is likely. However, robots remain largely black boxes, offering users little insight into their evolving capabilities and limitations. To address this gap, we present explainable object-oriented HRI (X-OOHRI), an augmented reality (AR) interface that conveys robot action possibilities and constraints through visual signifiers, radial menus, color coding, and explanation tags. Our system encodes object properties and robot limits into object-oriented structures using a vision-language model, allowing explanation generation on the fly and direct manipulation of virtual twins spatially aligned within a simulated environment. We integrate the end-to-end pipeline with a physical robot and showcase diverse use cases ranging from low-level pick-and-place to high-level instructions. Finally, we evaluate X-OOHRI through a user study and find that participants effectively issue object-oriented commands, develop accurate mental models of robot limitations, and engage in mixed-initiative resolution.

AR交互人机协作可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。