arXiv:2507.06701cs.LG2025-07

让机器通过无动作示范自我改进,实现大规模模仿学习

Value from Observations: Towards Large-Scale Imitation Learning via Self-Improvement

  • 用价值函数在专家与非专家数据间传递信息,支持无动作示范学习
  • 实验证明该方法在复杂数据分布下仍能有效提升性能
  • 适合希望构建可迭代优化的实用化模仿学习系统的研究者

模仿学习从观察(IfO)提供了一种大规模学习行为的强大方式:与行为克隆或离线强化学习不同,IfO可利用无动作示范,从而避免高成本的动作标注示范或奖励函数。然而,当前IfO研究多集中于理想化的双模态数据分布场景,限制了结果的实际意义。本文研究更复杂的分布,并提出一种新方法,使模型能从此类数据中学习,推动模仿学习向通过自我改进实现迭代演进的方向发展。该方法将基于强化学习的模仿学习适配至无动作示范,利用价值函数在专家与非专家数据间传递信息。通过全面评估,我们厘清了不同数据分布与算法适用性的关系,揭示了现有方法的局限性。研究结果为开发更鲁棒、更实用的IfO技术提供了关键洞见,助力实现可扩展的行为学习。

原文摘要 · Abstract (English)

Imitation Learning from Observation (IfO) offers a powerful way to learn behaviors at large-scale: Unlike behavior cloning or offline reinforcement learning, IfO can leverage action-free demonstrations and thus circumvents the need for costly action-labeled demonstrations or reward functions. However, current IfO research focuses on idealized scenarios with mostly bimodal-quality data distributions, restricting the meaningfulness of the results. In contrast, this paper investigates more nuanced distributions and introduces a method to learn from such data, moving closer to a paradigm in which imitation learning can be performed iteratively via self-improvement. Our method adapts RL-based imitation learning to action-free demonstrations, using a value function to transfer information between expert and non-expert data. Through comprehensive evaluation, we delineate the relation between different data distributions and the applicability of algorithms and highlight the limitations of established methods. Our findings provide valuable insights for developing more robust and practical IfO techniques on a path to scalable behaviour learning.

模仿学习自改进无动作示范强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。