arXiv:2603.28732cs.ROcs.CV2026-03被引 2

用人类第一视角数据构建可动物体3D场景图,提升机器人操作能力

Pandora: Articulated 3D Scene Graphs from Egocentric Vision

  • 利用人类佩戴Aria眼镜的视角数据,推断可动物体结构
  • 恢复的物体部件模型质量媲美先进方法,无需复杂传感器
  • 生成的3D场景图显著提升机器人执行隐藏物品取物任务的能力

当前机器人映射系统依赖自身传感器构建度量-语义场景表示,但受限于本体和技能,常无法探索如抽屉、高柜等区域。本文通过人类自然探索场景时佩戴Project Aria眼镜获取的第一视角数据,将人类对物体可动性的知识迁移至机器人。仅用简单启发式方法,即可从该数据中恢复可动部件模型,其质量与基于其他模态的顶尖方法相当。进一步,将这些模型整合进3D场景图,增强对物体动态及容器关系的理解。最终在波士顿动力Spot机器人上验证:仅凭3D场景图输入,即可成功完成隐藏目标物品的检索任务。

原文摘要 · Abstract (English)

Robotic mapping systems typically approach building metric-semantic scene representations from the robot's own sensors and cameras. However, these "first person" maps inherit the robot's own limitations due to its embodiment or skillset, which may leave many aspects of the environment unexplored. For example, the robot might not be able to open drawers or access wall cabinets. In this sense, the map representation is not as complete, and requires a more capable robot to fill in the gaps. We narrow these blind spots in current methods by leveraging egocentric data captured as a human naturally explores a scene wearing Project Aria glasses, giving a way to directly transfer knowledge about articulation from the human to any deployable robot. We demonstrate that, by using simple heuristics, we can leverage egocentric data to recover models of articulate object parts, with quality comparable to those of state-of-the-art methods based on other input modalities. We also show how to integrate these models into 3D scene graph representations, leading to a better understanding of object dynamics and object-container relationships. We finally demonstrate that these articulated 3D scene graphs enhance a robot's ability to perform mobile manipulation tasks, showcasing an application where a Boston Dynamics Spot is tasked with retrieving concealed target items, given only the 3D scene graph as input.

3D场景图可动物体第一人称视觉机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。