arXiv:2605.17017cs.LGcs.AI2026-05

让模仿学习在环境变化时仍稳定有效,仅靠离线数据就实现鲁棒性。

When Dynamics Shift, Robust Task Inference Wins: Offline Imitation Learning with Behavior Foundation Models Revisited

论文配图:When Dynamics Shift, Robust Task Inference Wins: Offline Imitation Learning with Behavior Foundation Models Revisited
图 1 · 摘自论文原文
  • 将任务推断建模为对抗性优化,应对动态变化的环境扰动。
  • 在摩擦、执行等变化下性能显著优于现有方法。
  • 无需修改预训练,适合实际动态场景中的部署需求。

行为基础模型(BFM)通过预训练无任务特性的表征,实现了可扩展的模仿学习(IL),但现有方法假设环境动态不变,难以应对真实世界中的摩擦、执行或传感器噪声变化。本文提出将BFM的任务推断重构为鲁棒最小-最大优化问题,可在不修改预训练的前提下适应最坏情况的动态扰动。据我们所知,这是首个仅依赖单一典型环境离线数据即实现动态变化鲁棒性的BFM框架。实验表明,该方法在动态扰动下显著优于标准BFM及鲁棒离线模仿学习基线,证明了鲁棒策略可在任务推断阶段完全实现,提升了BFM在动态环境中的实用性。

原文摘要 · Abstract (English)

Behavior Foundation Models (BFMs) enable scalable imitation learning (IL) by pretraining task-agnostic representations that can be rapidly adapted to new tasks. However, existing BFMs assume fixed environment dynamics, limiting their robustness under real-world shifts such as changes in friction, actuation, or sensor noise. We address this by formulating BFM task-inference as a robust minimax optimization problem, enabling adaptation to worst-case dynamics perturbations without modifying pretraining. To the best of our knowledge, this is the first BFM-based framework that achieves robustness to dynamics shifts while relying solely on offline data from a single nominal environment. Our approach significantly outperforms standard BFM and robust offline IL baselines under dynamics shifts. These results demonstrate that robust policy can be achieved entirely at task-inference time, improving the practicality of BFMs in dynamic settings.

模仿学习鲁棒性离线强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。