arXiv:2601.21718cs.LGcs.AI2026-01被引 3

预测逆动力学模型在数据少时更高效,能显著减少模仿学习所需演示次数。

When Does Predictive Inverse Dynamics Outperform Behavior Cloning?

  • 用未来状态预测来降低逆动力学模型的方差,提升稳定性。
  • 在2D任务中,行为克隆需3-5倍更多演示才能达到相同效果。
  • 适用于低样本场景,尤其适合高维视觉输入的复杂环境。

行为克隆(BC)是一种实用的离线模仿学习方法,但在专家示范数据有限时表现不佳。近期提出的预测逆动力学模型(PIDM)结合了未来状态预测器与逆动力学模型。尽管PIDM通常优于BC,其原因尚不明确。本文提供理论解释:PIDM存在权衡——基于预测的未来状态条件化可大幅降低方差,但预测本身引入额外偏差和方差。我们建立了PIDM相比BC实现更高样本效率和更低预测误差的条件,且在有额外数据源时优势更明显。我们在2D导航任务中验证了理论结果,发现BC需最多五倍(平均三倍)更多演示才能达到与PIDM相当的性能。在包含高维视觉输入和随机转换的现代视频游戏3D复杂环境中,BC所需样本超过PIDM的66%以上。

原文摘要 · Abstract (English)

Behavior cloning (BC) is a practical offline imitation learning method, but it often fails when expert demonstrations are limited. Recent works have introduced a class of architectures named predictive inverse dynamics models (PIDMs) that combine a future-state predictor with an inverse dynamics model. While PIDMs often outperform BC, the reasons behind their benefits remain unclear. In this paper, we provide a theoretical explanation: PIDMs introduce a tradeoff. Conditioning the IDM on the predicted future state can significantly reduce variance, but the prediction itself introduces additional bias and variance. We establish conditions for PIDMs to achieve higher sample efficiency and lower prediction error than BC, with the gap widening when additional data sources are available. We validate the theoretical insights empirically in 2D navigation tasks, where BC requires up to five times (three times on average) more demonstrations than PIDM to reach comparable performance. Results are also illustrated in a complex 3D environment in a modern video game with high-dimensional visual inputs and stochastic transitions, where BC requires over 66\% more samples than PIDM.

模仿学习逆动力学样本效率强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。