通过智能选目标,让机器人从示范中更高效地学习复杂任务。
Learning from Demonstrations via Capability-Aware Goal Sampling
- 根据智能体能力动态选取挑战性目标,构建自适应学习路径。
- 在稀疏奖励任务中,样本效率和最终成功率显著优于现有方法。
- 适合需要长期规划的机器人控制、游戏智能体等场景。
尽管模仿学习前景广阔,但在长时序环境中,完美复现示范不现实,微小误差会累积导致失败。本文提出Cago(能力感知目标采样),一种新型的从示范中学习的方法,缓解对专家轨迹的脆弱依赖。与以往仅用示范进行策略初始化或奖励设计的方法不同,Cago动态追踪智能体在专家轨迹上的能力水平,并利用该信号选择当前能力稍远的中间步骤作为目标,引导学习。这种方法形成自适应课程,推动智能体稳步完成完整任务。实验证明,Cago在多种稀疏奖励、目标条件化任务中显著提升样本效率和最终性能,持续优于现有模仿学习基线。
原文摘要 · Abstract (English)
Despite its promise, imitation learning often fails in long-horizon environments where perfect replication of demonstrations is unrealistic and small errors can accumulate catastrophically. We introduce Cago (Capability-Aware Goal Sampling), a novel learning-from-demonstrations method that mitigates the brittle dependence on expert trajectories for direct imitation. Unlike prior methods that rely on demonstrations only for policy initialization or reward shaping, Cago dynamically tracks the agent's competence along expert trajectories and uses this signal to select intermediate steps--goals that are just beyond the agent's current reach--to guide learning. This results in an adaptive curriculum that enables steady progress toward solving the full task. Empirical results demonstrate that Cago significantly improves sample efficiency and final performance across a range of sparse-reward, goal-conditioned tasks, consistently outperforming existing learning from-demonstrations baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。