arXiv:2409.11870cs.RO2024-09

让机器人通过互动学会识别开关并理解其功能

SpotLight: Robotic Scene Understanding through Interaction and Affordance Detection

  • 用视觉语言模型预测开关动作方式,指导真实交互
  • 实测操作成功率达84%,显著提升复杂任务执行能力
  • 支持机器人通过试错发现场景新关系,适合家用机器人研究

尽管家庭机器人研究不断深入,但部署于家庭环境的机器人在处理抽屉、灯开关等复杂功能部件时仍面临挑战,主要源于任务特定理解与交互能力不足。这些任务不仅需要检测与位姿估计,还需理解物体提供的使用可能性(即功能)。为应对这一问题,我们提出SpotLight:一个面向功能部件(特别是灯开关)交互的综合性框架,并使机器人可通过交互过程持续改进环境认知。该框架利用基于视觉语言模型的可操作性预测,估计灯开关的运动原型,在真实世界实验中实现最高84%的操作成功率。我们还构建了包含715张图像的专用数据集,并开发了针对灯开关检测的定制化模型。实验表明,该框架能通过物理互动促进机器人学习,使其在场景图表示中发现此前未知的关联。最后,我们拓展框架以支持如推拉门等其他功能性交互,展示其灵活性。视频与代码见:timengelbracht.github.io/SpotLight/

原文摘要 · Abstract (English)

Despite increasing research efforts on household robotics, robots intended for deployment in domestic settings still struggle with more complex tasks such as interacting with functional elements like drawers or light switches, largely due to limited task-specific understanding and interaction capabilities. These tasks require not only detection and pose estimation but also an understanding of the affordances these elements provide. To address these challenges and enhance robotic scene understanding, we introduce SpotLight: A comprehensive framework for robotic interaction with functional elements, specifically light switches. Furthermore, this framework enables robots to improve their environmental understanding through interaction. Leveraging VLM-based affordance prediction to estimate motion primitives for light switch interaction, we achieve up to 84% operation success in real world experiments. We further introduce a specialized dataset containing 715 images as well as a custom detection model for light switch detection. We demonstrate how the framework can facilitate robot learning through physical interaction by having the robot explore the environment and discover previously unknown relationships in a scene graph representation. Lastly, we propose an extension to the framework to accommodate other functional interactions such as swing doors, showcasing its flexibility. Videos and Code: timengelbracht.github.io/SpotLight/

机器人交互场景理解功能识别视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。