arXiv:2511.07942cs.LGcs.AI2025-11被引 1

让机器人在环境变化时仍能稳定模仿专家行为。

Balance Equation-based Distributionally Robust Offline Imitation Learning

  • 基于平衡方程构建鲁棒优化框架,仅用正常数据学策略
  • 在动态扰动下性能优于现有离线模仿学习方法
  • 适合部署环境与训练有差异的机器人控制场景

模仿学习在无法手动设计奖励函数或控制器的机器人和控制任务中表现优异。然而,标准模仿学习隐含假设训练与部署时环境动态保持不变,而现实中建模误差、参数变化和对抗扰动常导致转移动态漂移,造成性能严重下降。本文提出基于平衡方程的分布鲁棒离线模仿学习框架,仅利用正常动态下的专家演示数据,在不需额外环境交互的情况下学习鲁棒策略。将问题建模为过渡模型不确定集上的分布鲁棒优化,寻找在最坏情况过渡分布下最小化模仿损失的策略。关键创新在于,该鲁棒目标可完全转化为名义数据分布的形式,实现可计算的离线学习。在连续控制基准测试中,该方法在扰动或动态偏移环境下均表现出更优的鲁棒性和泛化能力,显著优于当前最优离线模仿学习基线。

原文摘要 · Abstract (English)

Imitation Learning (IL) has proven highly effective for robotic and control tasks where manually designing reward functions or explicit controllers is infeasible. However, standard IL methods implicitly assume that the environment dynamics remain fixed between training and deployment. In practice, this assumption rarely holds where modeling inaccuracies, real-world parameter variations, and adversarial perturbations can all induce shifts in transition dynamics, leading to severe performance degradation. We address this challenge through Balance Equation-based Distributionally Robust Offline Imitation Learning, a framework that learns robust policies solely from expert demonstrations collected under nominal dynamics, without requiring further environment interaction. We formulate the problem as a distributionally robust optimization over an uncertainty set of transition models, seeking a policy that minimizes the imitation loss under the worst-case transition distribution. Importantly, we show that this robust objective can be reformulated entirely in terms of the nominal data distribution, enabling tractable offline learning. Empirical evaluations on continuous-control benchmarks demonstrate that our approach achieves superior robustness and generalization compared to state-of-the-art offline IL baselines, particularly under perturbed or shifted environments.

模仿学习鲁棒控制离线学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。