arXiv:2410.12124cs.ROcs.AI2024-10中稿 · CoRL被引 12

仅用10次示范实现高效、可泛化的机器人操作策略

Learning from 10 Demos: Generalisable and Sample-Efficient Policy Learning with Oriented Affordance Frames

  • 用定向属性帧构建状态与动作的结构化表示
  • 10次示范即达成高成功率,泛化至未见物体与布局
  • 支持子策略组合,适合复杂多步骤任务

模仿学习已使机器人展现出高灵巧行为,但在长时序、多物体任务中仍受限于样本效率低和泛化能力差。现有方法需大量示范覆盖任务变化,成本高且难用于真实场景。本文提出定向属性帧,一种对状态与动作空间的结构化表示,显著提升空间与类别内泛化能力,使策略仅需10次示范即可高效学习。更重要的是,该抽象支持独立训练的子策略进行组合,解决长时序多物体任务。为实现子策略间平滑切换,我们引入自进展预测,直接从训练示范时长中推导。在三个真实世界任务上验证,尽管数据量小,策略仍能鲁棒泛化至未见物体外观、几何形状与空间排列,无需穷尽训练数据即可达到高成功率。视频演示见项目页:https://affordance-policy.github.io/。

原文摘要 · Abstract (English)

Imitation learning has unlocked the potential for robots to exhibit highly dexterous behaviours. However, it still struggles with long-horizon, multi-object tasks due to poor sample efficiency and limited generalisation. Existing methods require a substantial number of demonstrations to cover possible task variations, making them costly and often impractical for real-world deployment. We address this challenge by introducing oriented affordance frames, a structured representation for state and action spaces that improves spatial and intra-category generalisation and enables policies to be learned efficiently from only 10 demonstrations. More importantly, we show how this abstraction allows for compositional generalisation of independently trained sub-policies to solve long-horizon, multi-object tasks. To seamlessly transition between sub-policies, we introduce the notion of self-progress prediction, which we directly derive from the duration of the training demonstrations. We validate our method across three real-world tasks, each requiring multi-step, multi-object interactions. Despite the small dataset, our policies generalise robustly to unseen object appearances, geometries, and spatial arrangements, achieving high success rates without reliance on exhaustive training data. Video demonstration can be found on our project page: https://affordance-policy.github.io/.

模仿学习样本高效泛化能力机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。