用多步预测提升模仿学习的鲁棒性,减少误差累积。
A Model-Based Approach to Imitation Learning through Multi-Step Predictions
- 基于模型预测控制思想,通过多步状态预测改进模仿学习
- 在多种分布偏移和噪声下表现更优,超越传统行为克隆
- 提供理论保证,适合需高可靠性的决策任务
模仿学习是训练智能体复现专家行为的常用方法,但在复杂决策任务中常因误差累积和训练部署分布差异导致性能下降。本文提出一种受模型预测控制启发的新型模型基模仿学习框架,通过多步状态预测整合预测建模,有效缓解上述问题。实验表明,该方法在数值基准测试中优于传统行为克隆,在数据与执行过程中的分布偏移和测量噪声下均展现更强鲁棒性。此外,本文还提供了样本复杂度与误差界的理论保证,揭示了方法的收敛特性。
原文摘要 · Abstract (English)
Imitation learning is a widely used approach for training agents to replicate expert behavior in complex decision-making tasks. However, existing methods often struggle with compounding errors and limited generalization, due to the inherent challenge of error correction and the distribution shift between training and deployment. In this paper, we present a novel model-based imitation learning framework inspired by model predictive control, which addresses these limitations by integrating predictive modeling through multi-step state predictions. Our method outperforms traditional behavior cloning numerical benchmarks, demonstrating superior robustness to distribution shift and measurement noise both in available data and during execution. Furthermore, we provide theoretical guarantees on the sample complexity and error bounds of our method, offering insights into its convergence properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。