arXiv:2506.00083cs.ROcs.AI2025-06被引 2

构建分层动态场景图,让机器人在人机环境中自主完成复杂任务。

Hi-Dyna Graph: Hierarchical Dynamic Scene Graph for Robotic Autonomy in Human-Centric Environments

  • 分层结构融合全局拓扑与局部动态关系,支持环境演变下的实时更新。
  • 实测证明可在咖啡厅等动态场景中无需训练即完成任务,成功率高。
  • 适合需要上下文感知的移动操作机器人研究与应用开发者。

服务机器人在人机交互场景中的自主运行仍具挑战,需理解动态环境并做出上下文感知决策。现有方法如拓扑地图虽能提供高效空间先验,却难以建模临时物体关系;而密集神经表示(如NeRF)则计算成本过高。受分层场景表示与视频场景图生成启发,本文提出Hi-Dyna Graph,一种结合持久全局布局与局部动态语义的分层动态场景图架构,用于具身机器人自主。该框架从带姿态的RGB-D输入构建全局拓扑图,编码房间尺度连通性及大型静态物体(如家具);同时,环境与本体摄像头生成包含物体位置关系和人-物交互模式的动态子图。通过语义与空间约束将子图锚定于全局拓扑,实现环境演化下的无缝更新。采用大语言模型(LLM)驱动的智能体解析统一图,推断潜在任务触发条件,并生成基于机器人功能性的可执行指令。复杂实验验证了其优越的场景表征能力。真实部署测试表明,移动操作机器人在无额外训练或复杂奖励机制下,即可在动态场景中自主完成复杂任务,如担任咖啡厅助理。更多细节见https://anonymous.4open.science/r/Hi-Dyna-Graph-B326。

原文摘要 · Abstract (English)

Autonomous operation of service robotics in human-centric scenes remains challenging due to the need for understanding of changing environments and context-aware decision-making. While existing approaches like topological maps offer efficient spatial priors, they fail to model transient object relationships, whereas dense neural representations (e.g., NeRF) incur prohibitive computational costs. Inspired by the hierarchical scene representation and video scene graph generation works, we propose Hi-Dyna Graph, a hierarchical dynamic scene graph architecture that integrates persistent global layouts with localized dynamic semantics for embodied robotic autonomy. Our framework constructs a global topological graph from posed RGB-D inputs, encoding room-scale connectivity and large static objects (e.g., furniture), while environmental and egocentric cameras populate dynamic subgraphs with object position relations and human-object interaction patterns. A hybrid architecture is conducted by anchoring these subgraphs to the global topology using semantic and spatial constraints, enabling seamless updates as the environment evolves. An agent powered by large language models (LLMs) is employed to interpret the unified graph, infer latent task triggers, and generate executable instructions grounded in robotic affordances. We conduct complex experiments to demonstrate Hi-Dyna Grap's superior scene representation effectiveness. Real-world deployments validate the system's practicality with a mobile manipulator: robotics autonomously complete complex tasks with no further training or complex rewarding in a dynamic scene as cafeteria assistant. See https://anonymous.4open.science/r/Hi-Dyna-Graph-B326 for video demonstration and more details.

场景图机器人动态环境大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。