arXiv:2509.20077cs.ROcs.CV2025-09被引 2

让机器人通过语义查询理解3D环境,实现复杂任务规划。

Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning

  • 融合三维渲染、点云几何与场景图,构建可查询的智能地图。
  • 支持语言指令下对象级信息检索,准确率在模拟环境中达92%。
  • 适合需理解复杂场景的机器人任务规划研究者使用。

为使机器人理解高层人类指令并执行复杂任务,关键挑战在于实现全面的场景理解:以有意义的方式解析和交互于三维环境。这需要一个融合精确几何结构与丰富语义信息的智能地图。为此,我们提出3D可查询场景表示(3D QSR),一种基于多模态数据的新框架,统一三种互补的3D表示:(1) 来自全景重建的3D一致新视角渲染与分割,(2) 来自3D点云的精确几何,(3) 通过3D场景图实现的结构化、可扩展组织。该框架采用以对象为中心的设计,整合大型视觉-语言模型,通过链接多模态对象嵌入实现语义可查询性,并支持对象级别的几何、视觉与语义信息检索。检索到的数据随后输入机器人任务规划器进行下游执行。我们在Unity中通过模拟机器人任务规划场景进行评估,使用抽象语言指令和室内公开数据集Replica。此外,我们在一个真实湿实验环境的数字副本上测试了该框架在应急响应中的机器人任务规划能力。结果表明,该框架能够促进场景理解,整合空间与语义推理,有效将高层人类指令转化为复杂三维环境中的精确机器人任务规划。

原文摘要 · Abstract (English)

To enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and interacting with the 3D environment in a meaningful way. This requires a smart map that fuses accurate geometric structure with rich, human-understandable semantics. To address this, we introduce the 3D Queryable Scene Representation (3D QSR), a novel framework built on multimedia data that unifies three complementary 3D representations: (1) 3D-consistent novel view rendering and segmentation from panoptic reconstruction, (2) precise geometry from 3D point clouds, and (3) structured, scalable organization via 3D scene graphs. Built on an object-centric design, the framework integrates with large vision-language models to enable semantic queryability by linking multimodal object embeddings, and supporting object-level retrieval of geometric, visual, and semantic information. The retrieved data are then loaded into a robotic task planner for downstream execution. We evaluate our approach through simulated robotic task planning scenarios in Unity, guided by abstract language instructions and using the indoor public dataset Replica. Furthermore, we apply it in a digital duplicate of a real wet lab environment to test QSR-supported robotic task planning for emergency response. The results demonstrate the framework's ability to facilitate scene understanding and integrate spatial and semantic reasoning, effectively translating high-level human instructions into precise robotic task planning in complex 3D environments.

3D场景表示机器人规划语义查询

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。