用场景图增强机器人模仿学习的时空上下文理解能力
Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene Graphs

- 引入动态场景图作为结构化记忆,跟踪物体间关系随时间演变
- 在部分可观测环境下,长时序任务成功率提升显著
- 适合需要长期推理与环境适应的复杂家务机器人任务
模仿学习使机器人通过观察学习执行任务。然而,家庭和办公室等真实环境因空间尺度大,常存在严重部分观测问题。同时,许多任务需执行一系列子任务,要求机器人具备长时程推理能力。为此,我们提出在模仿学习中使用场景图作为显式且结构化的记忆机制。通过维护一个动态场景图,捕捉以物体为中心的关系及其随时间的变化,该方法使智能体在任务执行过程中保留相关历史上下文,从而高效推理逐步积累的场景信息。在模拟移动操作和真实桌面操作任务上的实验表明,该方法显著提升了策略性能,尤其在需要长时序推理和部分可观测性下具备强泛化能力。
原文摘要 · Abstract (English)
Imitation learning enables robots to learn how to execute tasks via observation. However, real-world environments like homes and offices are often severely partially observed due to their large spatial scales. In addition, many tasks involve executing a series of subtasks requiring autonomous robots to reason over extended time horizons. To address these challenges, we propose using scene graphs as an explicit and structured memory mechanism in imitation learning. By maintaining a dynamic scene graph that captures object-centric relationships and their evolution over time, our method allows the agent to retain relevant historical context during task execution to efficiently reason over incrementally accrued scene information. Our experiments on simulated mobile manipulation and real-world tabletop manipulation demonstrate that our approach substantially improves policy performance, particularly in settings that demand long-term reasoning and robust generalization under partial observability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。