用预训练动态表征,让少样本模仿学习更有效。
Offline Imitation Learning upon Arbitrary Demonstrations by Pre-Training Dynamics Representations
- 通过分解转移动态学习通用动态表征,减少下游学习参数。
- 仅需单条专家轨迹即可模仿出接近专家策略的行为。
- 适合数据稀缺场景,尤其适用于真实机器人部署。
有限的专家数据已成为制约离线模仿学习(Offline Imitation Learning, IL)扩展的关键瓶颈。本文提出在预训练阶段学习由转移动态分解得到的动态表征,以提升小样本下的IL性能。理论上证明,离线IL的最优决策变量位于该表征空间,显著降低下游学习所需参数量。动态表征可从具有相同动态的任意数据中学习,实现大量非专家数据的复用,缓解数据稀缺问题。我们设计了一种基于噪声对比估计的可计算损失函数用于预训练。在MuJoCo上的实验表明,所提算法仅需一条专家轨迹即可逼近专家策略;在真实四足机器人上,利用仿真数据预训练的动态表征,成功从少量真实演示中学会行走。
原文摘要 · Abstract (English)
Limited data has become a major bottleneck in scaling up offline imitation learning (IL). In this paper, we propose enhancing IL performance under limited expert data by introducing a pre-training stage that learns dynamics representations, derived from factorizations of the transition dynamics. We first theoretically justify that the optimal decision variable of offline IL lies in the representation space, significantly reducing the parameters to learn in the downstream IL. Moreover, the dynamics representations can be learned from arbitrary data collected with the same dynamics, allowing the reuse of massive non-expert data and mitigating the limited data issues. We present a tractable loss function inspired by noise contrastive estimation to learn the dynamics representations at the pre-training stage. Experiments on MuJoCo demonstrate that our proposed algorithm can mimic expert policies with as few as a single trajectory. Experiments on real quadrupeds show that we can leverage pre-trained dynamics representations from simulator data to learn to walk from a few real-world demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。