提出CubeDAgger,让机器人在动态任务中高效低风险地学习专家动作。
CubeDAgger: Interactive Imitation Learning for Dynamic Systems with Efficient yet Low-risk Interaction
- 用阈值正则化与多候选动作共识系统,优化专家干预时机。
- 仿真显示策略鲁棒性强,动态稳定性显著提升。
- 实机实验仅需30分钟交互,适合人机协作的机器人训练。
交互式模仿学习通过逐步专家指导提升智能体控制策略的鲁棒性。现有方法多采用专家-智能体切换机制以减少专家负担,但仅适用于静态任务;在动态任务中,干预时机偏差会导致动作突变,破坏机器人动态稳定性。本文提出新方法CubeDAgger,基于EnsembleDAgger进行三方面改进:第一,引入正则化显式激活决策阈值;第二,将切换系统重构为多动作候选的最优共识机制;第三,注入自回归彩色噪声以实现时间一致性的探索。仿真验证表明,所学策略兼具高鲁棒性与良好动态稳定性。真实机器人舀取实验(人类专家参与)进一步证明,仅需30分钟交互即可从零学习到稳健策略。
原文摘要 · Abstract (English)
Interactive imitation learning makes an agent's control policy robust by stepwise supervisions from an expert. The recent algorithms mostly employ expert-agent switching systems to reduce the expert's burden by limitedly selecting the supervision timing. However, this approach is useful only for static tasks; in dynamic tasks, timing discrepancies cause abrupt changes in actions, losing the robot's dynamic stability. This paper therefore proposes a novel method, named CubeDAgger, which improves robustness with less dynamic stability violations even for dynamic tasks. The proposed method is designed on a baseline, EnsembleDAgger, with three improvements. The first adds a regularization to explicitly activate the threshold for deciding the supervision timing. The second transforms the expert-agent switching system to an optimal consensus system of multiple action candidates. Third, autoregressive colored noise is injected to the agent's actions for time-consistent exploration. These improvements are verified by simulations, showing that the trained policies are sufficiently robust while maintaining dynamic stability during interaction. Finally, real-robot scooping experiments with a human expert demonstrate that the proposed method can learn robust policies from scratch based on just 30 minutes of interaction. https://youtu.be/kBl3SCTnVEM
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。