arXiv:2507.21533cs.ROcs.AI2025-07中稿 · ICLR被引 4

用观察数据训练可规划的智能体,提升模仿学习的效率与鲁棒性。

Model Predictive Adversarial Imitation Learning for Planning from Observation

  • 用规划代理替代策略,统一逆强化学习与模型预测控制
  • 少样本甚至单样本演示下仍实现高效规划,泛化能力更强
  • 适合需要安全、可解释规划的机器人任务

人类示范数据常存在模糊和不完整问题,促使人们发展具备可靠规划能力的模仿学习方法。一种常见范式是通过逆强化学习(IRL)学习奖励函数,再用模型预测控制(MPC)进行部署。本文提出将IRL中的策略替换为基于规划的代理,建立与对抗性模仿学习的联系,实现仅从观察数据中端到端学习规划器。在仿真控制基准和真实世界导航实验中,该方法在少量至单个观察示范条件下表现出显著更高的样本效率、分布外泛化能力与鲁棒性。

原文摘要 · Abstract (English)

Human demonstration data is often ambiguous and incomplete, motivating imitation learning approaches that also exhibit reliable planning behavior. A common paradigm to perform planning-from-demonstration involves learning a reward function via Inverse Reinforcement Learning (IRL) then deploying this reward via Model Predictive Control (MPC). Towards unifying these methods, we derive a replacement of the policy in IRL with a planning-based agent. With connections to Adversarial Imitation Learning, this formulation enables end-to-end interactive learning of planners from observation-only demonstrations. In addition to benefits in interpretability, complexity, and safety, we study and observe significant improvements on sample efficiency, out-of-distribution generalization, and robustness. The study includes evaluations in both simulated control benchmarks and real-world navigation experiments using few-to-single observation-only demonstrations.

模仿学习规划强化学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。