让机器人通过交互探索复杂环境,比纯视觉方法更智能。
CuriousBot: Interactive Mobile Exploration via Actionable 3D Relational Object Graph

- 构建3D关系物体图,让机器人理解物体间互动关系。
- 在多种场景中验证,探索效果优于仅依赖视觉语言模型的方法。
- 适合需要主动交互的移动机器人应用,如家庭服务、搜救。
移动探索是机器人领域的长期挑战,现有方法多聚焦于主动感知而非主动交互,限制了机器人与环境的深度互动能力。当前基于交互的探索方法通常局限于桌面场景,忽视了移动探索中的大空间、复杂动作空间和多样物体关系等独特挑战。本文提出一种3D关系物体图,用于编码丰富的物体间关系,并基于此构建可实现主动交互的探索系统。我们在多样化场景中评估该系统,定性和定量结果表明其在不同物体实例、关系和场景间均具有优异的泛化能力,性能超越仅依赖视觉语言模型(VLMs)的方法。
原文摘要 · Abstract (English)
Mobile exploration is a longstanding challenge in robotics, yet current methods primarily focus on active perception instead of active interaction, limiting the robot's ability to interact with and fully explore its environment. Existing robotic exploration approaches via active interaction are often restricted to tabletop scenes, neglecting the unique challenges posed by mobile exploration, such as large exploration spaces, complex action spaces, and diverse object relations. In this work, we introduce a 3D relational object graph that encodes diverse object relations and enables exploration through active interaction. We develop a system based on this representation and evaluate it across diverse scenes. Our qualitative and quantitative results demonstrate the system's effectiveness and generalization across object instances, relations, and scenes, outperforming methods solely relying on vision-language models (VLMs).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。