arXiv:2602.02762cs.LG2026-02中稿 · ICML被引 1

揭示逆动力学模型在半监督模仿学习中的高效性来源

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning

  • 用逆动力学模型生成动作标签,提升小样本模仿学习效率
  • 实验验证逆动力学模型比行为克隆更省数据,尤其在少量标注数据下
  • 适合研究半监督模仿学习、数据高效策略训练的科研人员

半监督模仿学习(SSIL)旨在仅用少量带动作标签的轨迹和大量无标签轨迹学习策略。部分方法通过逆动力学模型(IDM)从当前状态与下一状态预测动作,可作为策略使用(VM-IDM)或为无标签数据生成动作标签(IDM标注)。本文首先证明,在极限情况下,VM-IDM与IDM标注会收敛到同一策略,称为基于IDM的策略。我们进一步指出,此前观察到的该策略优于行为克隆的现象,源于IDM学习本身的高样本效率,原因有二:(i) 真实的IDM属于复杂度更低的假设空间;(ii) 真实的IDM通常比专家策略更少随机性。基于统计学习理论与新实验(包括使用统一视频-动作预测架构UVA的研究),我们提出改进的LAPO算法用于潜在动作策略学习。实验在Procgen、Push-T和LIBERO基准上验证了方法有效性。

原文摘要 · Abstract (English)

Semi-supervised imitation learning (SSIL) consists in learning a policy from a small dataset of action-labeled trajectories and a much larger dataset of action-free trajectories. Some SSIL methods learn an inverse dynamics model (IDM) to predict the action from the current state and the next state. An IDM can act as a policy when paired with a video model (VM-IDM) or as a label generator to perform behavior cloning on action-free data (IDM labeling). In this work, we first show that VM-IDM and IDM labeling learn the same policy in a limit case, which we call the IDM-based policy. We then argue that the previously observed advantage of IDM-based policies over behavior cloning is due to the superior sample efficiency of IDM learning, which we attribute to two causes: (i) the ground-truth IDM tends to be contained in a lower complexity hypothesis class relative to the expert policy, and (ii) the ground-truth IDM is often less stochastic than the expert policy. We argue these claims based on insights from statistical learning theory and novel experiments, including a study of IDM-based policies using recent architectures for unified video-action prediction (UVA). Motivated by these insights, we finally propose an improved version of the existing LAPO algorithm for latent action policy learning. We experiment on the Procgen, Push-T and LIBERO benchmarks.

模仿学习样本效率逆动力学半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。