arXiv:2503.13162cs.LGcs.AI2025-03ICLR被引 4

在机器人模仿人类行为时出错的情况下,如何少试错却学得更好。

Efficient Imitation under Misspecification

  • 提出新条件确保局部搜索不放大错误。
  • 扩大搜索范围至学习者能实际执行的策略可达状态。
  • 适合研究模仿学习鲁棒性与安全性的学者。

我们研究模仿学习中的模型误设问题:学习者在某些情况下无法完全复现专家行为,这通常源于观测空间或动作空间表达能力的差异(如机器人与人类的感知或形态差异)。由于学习者必然存在错误,必须通过与环境交互来识别哪些错误代价高并引发累积误差。但交互成本高且存在安全隐患,因此需尽可能减少交互次数同时保证学到强策略。已有工作提出高效的逆强化学习算法,在可实现条件下仅进行计算高效的局部搜索并有严格保障。本文首次证明,在一种新提出的奖励无关策略完备性条件下,这类基于局部搜索的逆强化学习算法能避免误差累积。接着探讨应在哪里开展局部搜索——在误设场景下,学习者可能不如专家擅长“走钢丝”。我们证明,在误设情况下,将局部搜索扩展至学习者可实际执行的良好策略所达状态是有益的。最后,通过实验探索多种误设来源,并研究离线数据如何有效拓宽局部搜索范围。

原文摘要 · Abstract (English)

We consider the problem of imitation learning under misspecification: settings where the learner is fundamentally unable to replicate expert behavior everywhere. This is often true in practice due to differences in observation space and action space expressiveness (e.g. perceptual or morphological differences between robots and humans). Given the learner must make some mistakes in the misspecified setting, interaction with the environment is fundamentally required to figure out which mistakes are particularly costly and lead to compounding errors. However, given the computational cost and safety concerns inherent in interaction, we'd like to perform as little of it as possible while ensuring we've learned a strong policy. Accordingly, prior work has proposed a flavor of efficient inverse reinforcement learning algorithms that merely perform a computationally efficient local search procedure with strong guarantees in the realizable setting. We first prove that under a novel structural condition we term reward-agnostic policy completeness, these sorts of local-search based IRL algorithms are able to avoid compounding errors. We then consider the question of where we should perform local search in the first place, given the learner may not be able to "walk on a tightrope" as well as the expert in the misspecified setting. We prove that in the misspecified setting, it is beneficial to broaden the set of states on which local search is performed to include those reachable by good policies the learner can actually play. We then experimentally explore a variety of sources of misspecification and how offline data can be used to effectively broaden where we perform local search from.

模仿学习逆强化学习鲁棒性策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。