用单目相机实时构建可动态调整细节的3D场景图,支持机器人边操作边学习新任务。
FOUND-IT: Foundation-model-first Task-driven 3D Scene Graphs with Granularity on Demand

- 基于几何基础模型,通过新增模块直接重建可通行区域信息。
- 在ASHiTA SG3D基准上准确率比现有方法高79%,支持实时运行。
- 适用于未预设任务的复杂人机协同场景,适合智能机器人应用。
我们提出首个使用未校准单目相机实时构建任意室内外环境层次化任务驱动3D场景图的方法。利用几何基础模型估计场景图中的几何属性(如物体边界框),同时发现可通过在现有几何基础模型(如VGGT)上添加额外头,直接重建可通行性信息(场景图的“地点”层)。该方法具有任务驱动特性:根据任务动态调整对象与区域的粒度——例如,在操作任务中能识别炉灶上的小旋钮,在导航任务中则关注整体大物体(如整台炉灶)。与以往工作不同,我们考虑任务列表不预先固定、随机器人运行而演变的真实场景,自然支持复杂的移动-操作任务,使机器人可随任务进展动态调整表征。我们称此方法为FOUND-IT。该系统还包含一种代理式查询机制以获取场景图信息。在ASHiTA SG3D任务定位基准上,准确率提升79%;我们还在搭载Jetson Thor的地面机器人上实现实时运行。此外,为验证鲁棒性,我们在YouTube随手拍摄的房地产公寓导览视频上成功构建了3D场景图。代码将在发表后公开。
原文摘要 · Abstract (English)
We present the first approach to build hierarchical task-driven 3D scene graphs of arbitrary indoor or outdoor environments using an uncalibrated monocular camera in real-time. We leverage geometric foundation models to estimate geometric attributes of the scene graph (e.g., object bounding boxes), but we also observe that traversability information (the "places" layer of a scene graph) can be directly reconstructed by adding an extra head to existing geometric foundation models, like VGGT. Our approach is task-driven in the sense that we adjust the granularity of the objects and regions in the map depending on the task; for instance, during a manipulation task, our approach is able to resolve small knobs on a stove, while during a navigation task it can focus on large objects (e.g., the entire stove). However, in a major departure from related work, we consider the realistic case where the list of tasks is not predefined and fixed, but evolves as the robot operates. This naturally allows dealing with complex loco-manipulation tasks, where the robot can dynamically adjust its representation as the task unfolds. We dub the resulting approach FOUND-IT. FOUND-IT also includes an agentic approach to query information in the scene graph. In addition to achieving 79% higher accuracy on the ASHiTA SG3D task grounding benchmark, we demonstrate FOUND-IT runs in real-time on a ground robot using a Jetson Thor. Furthermore, to highlight the robustness of our method, we demonstrate constructing 3D scene graphs on casually captured realtor apartment tours from YouTube. Code will be made available upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。