arXiv:2506.22593cs.ROcs.AI2025-06中稿 · 2025 IEEE Internat…被引 4

将图像与激光雷达实时转为结构化场景图,实现人机协同理解

Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding

  • 基于图像和激光雷达数据,纯在CPU上实时生成结构化场景图
  • 输出去噪2D地图与分割3D点云,支持从物体到建筑的多层抽象
  • 适用于资源受限机器人,在真实车库和办公环境实现自主探索

自主机器人在高风险场景中日益成为人类操作员的重要辅助平台。为完成复杂任务,需高效的人机协作与理解。虽然机器人规划依赖三维几何信息,但人类更习惯于高层级的二维地图,如表示建筑信息模型(BIM)的俯视图。三维场景图已成为连接人类可读的二维BIM与机器人三维地图的有力工具。本文提出像素转图(Pix2G),一种轻量级方法,可在资源受限的机器人平台上,仅使用CPU实时从图像像素和激光雷达地图生成结构化场景图。该方法输出去噪后的二维俯视环境图和结构分割的三维点云,并通过多层图结构无缝连接,涵盖从物体级到建筑级的信息。在NASA JPL NeBula-Spot足式机器人上进行的真实世界实验表明,该方法能实时自主探索并映射杂乱的车库及类办公室环境。

原文摘要 · Abstract (English)

Autonomous robots are increasingly playing key roles as support platforms for human operators in high-risk, dangerous applications. To accomplish challenging tasks, an efficient human-robot cooperation and understanding is required. While typically robotic planning leverages 3D geometric information, human operators are accustomed to a high-level compact representation of the environment, like top-down 2D maps representing the Building Information Model (BIM). 3D scene graphs have emerged as a powerful tool to bridge the gap between human readable 2D BIM and the robot 3D maps. In this work, we introduce Pixels-to-Graph (Pix2G), a novel lightweight method to generate structured scene graphs from image pixels and LiDAR maps in real-time for the autonomous exploration of unknown environments on resource-constrained robot platforms. To satisfy onboard compute constraints, the framework is designed to perform all operation on CPU only. The method output are a de-noised 2D top-down environment map and a structure-segmented 3D pointcloud which are seamlessly connected using a multi-layer graph abstracting information from object-level up to the building-level. The proposed method is quantitatively and qualitatively evaluated during real-world experiments performed using the NASA JPL NeBula-Spot legged robot to autonomously explore and map cluttered garage and urban office like environments in real-time.

场景图人机协作实时推理机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。