连续动作控制中,模仿学习会随时间指数放大误差。
The Pitfalls of Imitation Learning when Actions are Continuous
- 即使专家策略平滑确定,模仿者也会因动作连续性产生指数级执行误差。
- 在长时域任务中,纯从专家数据学习的算法必然导致误差指数增长。
- 突破瓶颈需非光滑、非马尔可夫或高状态依赖随机策略,如扩散模型。
我们研究离散时间、连续状态与动作控制系统的专家示范模仿问题。即使系统动态满足指数稳定性(扰动影响呈指数衰减),且专家策略平滑确定,任何平滑、确定的模仿策略在执行时的误差仍会随问题时长远大于训练数据分布下的误差,且呈指数级增长。该负结果适用于仅从专家数据学习的任何算法,包括行为克隆和离线强化学习,除非算法生成高度‘不合规’的模仿策略——即非光滑、非马尔可夫或具有高度状态依赖随机性的策略,或专家轨迹分布足够‘分散’。我们通过实验验证了这些复杂策略参数化的优势,解释了当前机器人学习中流行策略(如动作分块与扩散策略)的有效性。此外,我们还建立了控制中模仿学习的一系列互补的负结果与正结果。
原文摘要 · Abstract (English)
We study the problem of imitating an expert demonstrator in a discrete-time, continuous state-and-action control system. We show that, even if the dynamics satisfy a control-theoretic property called exponential stability (i.e. the effects of perturbations decay exponentially quickly), and the expert is smooth and deterministic, any smooth, deterministic imitator policy necessarily suffers error on execution that is exponentially larger, as a function of problem horizon, than the error under the distribution of expert training data. Our negative result applies to any algorithm which learns solely from expert data, including both behavior cloning and offline-RL algorithms, unless the algorithm produces highly "improper" imitator policies--those which are non-smooth, non-Markovian, or which exhibit highly state-dependent stochasticity--or unless the expert trajectory distribution is sufficiently "spread." We provide experimental evidence of the benefits of these more complex policy parameterizations, explicating the benefits of today's popular policy parameterizations in robot learning (e.g. action-chunking and diffusion policies). We also establish a host of complementary negative and positive results for imitation in control systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。