提出分层物体分解方法,让机器人更高效地学习视觉运动控制。
Task-Oriented Hierarchical Object Decomposition for Visuomotor Control
- 按场景实体构建分层表示家族,按任务选择性组合
- 在10个任务中实现更高效的模仿学习,优于现有方法
- 支持零样本技能链,适合复杂真实场景的机器人控制
良好的预训练视觉表征可使机器人高效学习视觉运动策略。然而,现有表征采用通用化方法,存在两大缺陷:(1) 完全与任务无关,无法有效忽略场景中的无关信息;(2) 缺乏处理无约束/复杂现实场景的表征能力。为此,我们提出一种面向任务的分层物体分解表示(HODOR),通过按场景实体(物体及其部件)组织大规模组合式表征家族,实现按任务选择性组装,并随场景与任务复杂度动态扩展表征容量。实验表明,HODOR在5个模拟和5个真实世界操作任务中均优于现有的场景向量表示和物体中心表示,在样本效率上表现更优。此外,HODOR捕捉的不变性可传递至下游策略,使其在分布外测试条件下具备鲁棒泛化能力,支持零样本技能链。附录、代码及视频见:https://sites.google.com/view/hodor-corl24。
原文摘要 · Abstract (English)
Good pre-trained visual representations could enable robots to learn visuomotor policy efficiently. Still, existing representations take a one-size-fits-all-tasks approach that comes with two important drawbacks: (1) Being completely task-agnostic, these representations cannot effectively ignore any task-irrelevant information in the scene, and (2) They often lack the representational capacity to handle unconstrained/complex real-world scenes. Instead, we propose to train a large combinatorial family of representations organized by scene entities: objects and object parts. This hierarchical object decomposition for task-oriented representations (HODOR) permits selectively assembling different representations specific to each task while scaling in representational capacity with the complexity of the scene and the task. In our experiments, we find that HODOR outperforms prior pre-trained representations, both scene vector representations and object-centric representations, for sample-efficient imitation learning across 5 simulated and 5 real-world manipulation tasks. We further find that the invariances captured in HODOR are inherited into downstream policies, which can robustly generalize to out-of-distribution test conditions, permitting zero-shot skill chaining. Appendix, code, and videos: https://sites.google.com/view/hodor-corl24.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。