通过动态课程与对比学习,提升机器人操作的强化学习效率。
ACDC: Adaptive Curriculum Planning with Dynamic Contrastive Control for Goal-Conditioned Reinforcement Learning in Robotic Manipulation
- 根据成功率与进度动态调整学习路径,平衡探索与利用。
- 在复杂任务中实现更高样本效率和最终成功率。
- 适合研究机器人强化学习与自适应训练策略的读者。
目标条件强化学习在机器人操作中展现出巨大潜力,但现有方法受限于对经验优先级的依赖,导致在多样化任务中表现不佳。受人类学习行为启发,我们提出更全面的学习范式ACDC,整合多维度自适应课程(AC)规划与动态对比(DC)控制,引导智能体沿优化的学习轨迹前进。具体而言,在规划层面,AC组件基于智能体的成功率与训练进度,动态平衡多样性驱动的探索与质量驱动的利用;在控制层面,DC组件通过范数约束的对比学习实施课程计划,实现与当前课程重点一致的幅度引导经验选择。大量实验表明,ACDC在挑战性机器人操作任务中持续优于最先进基线,在样本效率和最终任务成功率上均表现更优。
原文摘要 · Abstract (English)
Goal-conditioned reinforcement learning has shown considerable potential in robotic manipulation; however, existing approaches remain limited by their reliance on prioritizing collected experience, resulting in suboptimal performance across diverse tasks. Inspired by human learning behaviors, we propose a more comprehensive learning paradigm, ACDC, which integrates multidimensional Adaptive Curriculum (AC) Planning with Dynamic Contrastive (DC) Control to guide the agent along a well-designed learning trajectory. More specifically, at the planning level, the AC component schedules the learning curriculum by dynamically balancing diversity-driven exploration and quality-driven exploitation based on the agent's success rate and training progress. At the control level, the DC component implements the curriculum plan through norm-constrained contrastive learning, enabling magnitude-guided experience selection aligned with the current curriculum focus. Extensive experiments on challenging robotic manipulation tasks demonstrate that ACDC consistently outperforms the state-of-the-art baselines in both sample efficiency and final task success rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。