arXiv:2507.09061cs.LGcs.SY2025-07被引 18

通过动作分块和探索性数据收集,显著提升连续控制中的行为克隆性能。

Action Chunking and Exploratory Data Collection Yield Exponential Improvements in Behavior Cloning for Continuous Control

  • 采用动作分块与探索性数据增强,降低误差累积。
  • 实验验证在多个机器人基准上性能呈指数级提升。
  • 适合研究模仿学习与机器人控制的学者参考。

本文对现代机器人连续控制中两种最有效的示范学习干预措施——动作分块(开环预测动作序列)和探索性专家示范增强——进行了理论分析。尽管现有研究表明,模仿学习在连续设置下存在随任务时域增长而指数级累积的误差问题,本文证明动作分块和探索性数据采集可在不同场景下有效规避此类误差累积。结果表明,控制论稳定性是这些干预措施效果的核心机制。在实证层面,我们在主流机器人学习基准上验证了预测结果,并确认控制论视角对误差累积机制提供了细粒度理解;在理论层面,该视角带来了比仅依赖信息论方法更紧致的模仿学习误差统计保证。

原文摘要 · Abstract (English)

This paper presents a theoretical analysis of two of the most impactful interventions in modern learning from demonstration in robotics and continuous control: the practice of action-chunking (predicting sequences of actions in open-loop) and exploratory augmentation of expert demonstrations. Though recent results show that learning from demonstration, also known as imitation learning (IL), can suffer errors that compound exponentially with task horizon in continuous settings, we demonstrate that action chunking and exploratory data collection circumvent exponential compounding errors in different regimes. Our results identify control-theoretic stability as the key mechanism underlying the benefits of these interventions. On the empirical side, we validate our predictions and the role of control-theoretic stability through experimentation on popular robot learning benchmarks. On the theoretical side, we demonstrate that the control-theoretic lens provides fine-grained insights into how compounding error arises, leading to tighter statistical guarantees on imitation learning error when these interventions are applied than previous techniques based on information-theoretic considerations alone.

模仿学习机器人控制误差累积

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。