用场景图聚焦关键物体,提升机器人组合技能的鲁棒性。
Compose by Focus: Scene Graph-based Atomic Skills
- 用场景图过滤无关干扰,专注任务相关对象与关系。
- 在仿真和真实场景中,成功率显著高于现有方法。
- 适合需要长期复杂任务的通用机器人系统。
通用机器人的一项关键能力是组合泛化——将基础技能组合以完成复杂、长时序任务。以往工作多关注预学习技能的规划排序,但执行单个技能仍面临挑战,因视觉运动策略在场景组合引发的分布偏移下易失效。为此,本文提出基于场景图的表示,聚焦任务相关物体与关系,降低对无关变化的敏感性。在此基础上,构建融合图神经网络与扩散式模仿学习的场景图技能学习框架,并将‘聚焦’的场景图技能与基于视觉语言模型(VLM)的任务规划器结合。在仿真与真实机械臂操作任务中,实验显示其成功率显著优于当前最优基线,验证了在长时序任务中更强的鲁棒性与组合泛化能力。
原文摘要 · Abstract (English)
A key requirement for generalist robots is compositional generalization - the ability to combine atomic skills to solve complex, long-horizon tasks. While prior work has primarily focused on synthesizing a planner that sequences pre-learned skills, robust execution of the individual skills themselves remains challenging, as visuomotor policies often fail under distribution shifts induced by scene composition. To address this, we introduce a scene graph-based representation that focuses on task-relevant objects and relations, thereby mitigating sensitivity to irrelevant variation. Building on this idea, we develop a scene-graph skill learning framework that integrates graph neural networks with diffusion-based imitation learning, and further combine "focused" scene-graph skills with a vision-language model (VLM) based task planner. Experiments in both simulation and real-world manipulation tasks demonstrate substantially higher success rates than state-of-the-art baselines, highlighting improved robustness and compositional generalization in long-horizon tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。