通过人类操作演示构建可动物体的3D功能场景图,提升机器人对复杂物体的理解能力。
ArtiSG: Functional 3D Scene Graph Construction via Human-demonstrated Articulated Objects Manipulation
- 利用人体操作轨迹数据,结合便携式硬件追踪6自由度运动,推断物体运动轴。
- 在真实环境中实现92%的功能元素召回率,比基线提升显著。
- 适合需要理解可动物体结构的机器人任务,如语言指令下的物品操作。
3D场景图已使机器人具备导航与规划所需的语义理解能力。然而,现有功能场景图多聚焦于静态元素检测,缺乏物理操作所需的运动学信息,尤其在可动物体方面表现不足。现有从静态观察中推断运动机制的方法易受视觉模糊影响,而基于状态变化估计参数的方法通常依赖固定摄像头和无遮挡视角。此外,隐藏把手等隐蔽功能部件常被纯视觉感知遗漏。为此,我们提出ArtiSG框架,通过将人类操作示范编码为结构化机器人记忆,构建功能性3D场景图。该方法采用便携式硬件系统,即使在相机自运动情况下仍能精确追踪6-DoF操作轨迹并估计运动轴。通过将这些运动先验融入分层开放词汇图,系统不仅能建模可动物体的运动方式,还可利用物理交互数据发现隐含功能元素。大量真实环境实验表明,ArtiSG在功能元素召回率和运动轴估计精度上显著优于基线。此外,所构建图谱作为可靠机器人记忆,有效指导机器人在包含多种可动物体的真实环境中执行语言引导的操纵任务。
原文摘要 · Abstract (English)
3D scene graphs have empowered robots with semantic understanding for navigation and planning. However, current functional scene graphs primarily focus on static element detection, lacking the actionable kinematic information required for physical manipulation, particularly regarding articulated objects. Existing approaches for inferring articulation mechanisms from static observations are prone to visual ambiguity, while methods that estimate parameters from state changes typically rely on constrained settings such as fixed cameras and unobstructed views. Furthermore, inconspicuous functional elements like hidden handles are frequently missed by pure visual perception. To bridge this gap, we present ArtiSG, a framework that constructs functional 3D scene graphs by encoding human demonstrations into structured robotic memory. Our approach leverages a robust data collection pipeline utilizing a portable hardware setup to accurately track 6-DoF manipulation trajectories and estimate articulation axes, even under camera ego-motion. By integrating these kinematic priors into a hierarchical, open-vocabulary graph, our system not only models how articulated objects move but also utilizes physical interaction data to discover implicit elements. Extensive real-world experiments demonstrate that ArtiSG significantly outperforms baselines in functional element recall and articulation estimation precision. Moreover, we show that the constructed graph serves as a reliable robotic memory, effectively guiding robots to perform language-directed manipulation tasks in real-world environments containing diverse articulated objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。