arXiv:2411.02938cs.RO2024-11中稿 · the Workshop on Li…被引 1

让机器人实时更新3D场景图,应对动态环境变化。

Multi-Modal 3D Scene Graph Updater for Shared and Dynamic Environments

  • 融合人、机器人感知、时间与动作等多模态信息更新场景图
  • 实现在动态环境中保持场景图一致性,支持高阶推理与规划
  • 适合研究具身智能、机器人导航与动态环境建模的学者

通用大语言模型(LLMs)和大视觉模型(VLMs)的出现,简化了语义丰富地图的构建,使机器人能够将其表示与高层推理和规划相结合。目前最常用的语义地图格式是3D场景图,它同时包含度量(低层)和语义(高层)信息。然而,这些地图通常假设世界是静态的,而真实环境如家庭和办公室是动态的。即使微小的变化也可能显著影响任务表现。为使机器人融入动态环境,必须实时检测变化并更新场景图。该更新过程本质上是多模态的,需整合人类代理、机器人自身感知系统、时间以及其行为等多源输入。本文提出一个框架,利用这些多模态输入在实时运行中维持场景图的一致性,展示出有前景的初步结果,并勾勒出未来研究路线。

原文摘要 · Abstract (English)

The advent of generalist Large Language Models (LLMs) and Large Vision Models (VLMs) have streamlined the construction of semantically enriched maps that can enable robots to ground high-level reasoning and planning into their representations. One of the most widely used semantic map formats is the 3D Scene Graph, which captures both metric (low-level) and semantic (high-level) information. However, these maps often assume a static world, while real environments, like homes and offices, are dynamic. Even small changes in these spaces can significantly impact task performance. To integrate robots into dynamic environments, they must detect changes and update the scene graph in real-time. This update process is inherently multimodal, requiring input from various sources, such as human agents, the robot's own perception system, time, and its actions. This work proposes a framework that leverages these multimodal inputs to maintain the consistency of scene graphs during real-time operation, presenting promising initial results and outlining a roadmap for future research.

3D场景图多模态动态环境机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。