让机器人从高效动作中自我学习,提升长期任务的执行效率。
Temporal Self-Imitation Learning

- 用成功轨迹中的高效动作作为自我监督信号。
- 在15个任务中,训练效率和完成速度平均提升30%以上。
- 适合长期规划、需要高效执行的机器人控制场景。
使用奖励塑造训练的长时程机器人操作策略,仍可能通过低效交互获得高回报,而训练中偶然发现的高效行为可能被遗忘。我们提出时间自模仿学习(TSIL),一种利用学习过程中生成的高效成功轨迹作为可复用监督信号的强化学习框架。TSIL通过配置相关的自适应时间目标逐步优化策略,并通过效率加权的自模仿机制保留并重放高效行为。在15个不同长时程操作任务上,TSIL一致提升了学习效率、任务完成效率、对快速成功行为的重复访问率,并增强了对不稳定训练条件的鲁棒性。结果表明,成功行为的时间结构本身即可提供超越人工奖励塑造的可扩展自监督信号。
原文摘要 · Abstract (English)
Long-horizon robot manipulation policies trained with reward shaping can still achieve high return through inefficient interactions, while rare efficient behaviors discovered during training may be forgotten. We argue that temporal efficiency itself provides a powerful and underutilized source of self-supervision for reinforcement learning. We introduce Temporal Self-Imitation Learning (TSIL), a reinforcement learning framework that mines temporally efficient successful trajectories generated during learning and converts them into reusable supervision for future policy improvement. TSIL progressively refines learning using configuration-conditioned adaptive temporal targets derived from fast successful trajectories, while preserving and replaying efficient behaviors through efficiency-weighted self-imitation learning. Across 15 distinct long-horizon manipulation tasks, TSIL consistently improves learning efficiency, task-completion efficiency, revisitation of fast successful behaviors, and robustness to unstable training conditions. More broadly, our results suggest that the temporal structure of successful behavior itself provides a scalable self-supervisory signal for reinforcement learning beyond manually engineered reward shaping alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。