用观察数据学动作表示,大幅减少机器人模仿所需标注数据。
Imitation Learning with Limited Actions via Diffusion Planners and Deep Koopman Controllers
- 先用扩散规划生成动作序列,再通过深度库普曼模型学习隐空间动作表征。
- 仅需少量带标签动作数据,就能实现高成功率的机器人任务执行。
- 适合数据稀缺场景,尤其适用于真实机器人上多模态行为模仿。
基于扩散的机器人策略虽在模仿多模态行为方面展现潜力,但通常需要大量带标签动作的示范数据,带来沉重的数据采集负担。本文提出一种‘规划-控制’框架,通过利用仅含观测数据的示范轨迹,提升逆动力学控制器的动作数据效率。具体地,采用深度库普曼算子建模系统动态,并基于纯观测轨迹学习隐空间动作表示。该表示可经由线性动作解码器映射为真实的高维连续动作,仅需极少动作标注数据。在模拟机器人操作任务及真实机器人上的多模态专家示范实验中,验证了该方法显著提升动作数据效率,并在有限动作数据下实现高任务成功率。
原文摘要 · Abstract (English)
Recent advances in diffusion-based robot policies have demonstrated significant potential in imitating multi-modal behaviors. However, these approaches typically require large quantities of demonstration data paired with corresponding robot action labels, creating a substantial data collection burden. In this work, we propose a plan-then-control framework aimed at improving the action-data efficiency of inverse dynamics controllers by leveraging observational demonstration data. Specifically, we adopt a Deep Koopman Operator framework to model the dynamical system and utilize observation-only trajectories to learn a latent action representation. This latent representation can then be effectively mapped to real high-dimensional continuous actions using a linear action decoder, requiring minimal action-labeled data. Through experiments on simulated robot manipulation tasks and a real robot experiment with multi-modal expert demonstrations, we demonstrate that our approach significantly enhances action-data efficiency and achieves high task success rates with limited action data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。