用时间方差指导学习,让智能体更高效掌握多目标任务。
TEACH: Temporal Variance-Driven Curriculum for Reinforcement Learning
- 根据策略信心波动大小动态选难目标,引导学习重点。
- 在11个机器人任务中显著提升训练效率,超越现有方法。
- 适用于各类强化学习框架,尤其适合复杂多目标场景。
强化学习在单目标任务中已取得显著成果,但在多目标设置下,统一的目标选择常导致样本效率低下,因为智能体需学习通用的目标条件策略。受生物系统自适应、结构化学习过程的启发,我们提出一种基于时间方差驱动课程的师生学习范式(TEACH),以加速目标条件强化学习。在此框架中,教师模块通过状态-动作价值函数(Q)参数化,动态优先选择策略置信度时间方差最高的目标。教师通过聚焦这些高不确定性目标,提供自适应且集中的学习信号,推动持续高效的进展。我们建立了Q值时间方差与策略演化的理论联系,揭示了该方法的内在机制。本方法具备算法无关性,可无缝集成至现有强化学习框架。我们在11个多样化的机器人操作与迷宫导航任务上进行了评估,结果表明其在一致性与显著性上均优于当前最先进的课程学习和目标选择方法。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has achieved significant success in solving single-goal tasks. However, uniform goal selection often results in sample inefficiency in multi-goal settings where agents must learn a universal goal-conditioned policy. Inspired by the adaptive and structured learning processes observed in biological systems, we propose a novel Student-Teacher learning paradigm with a Temporal Variance-Driven Curriculum to accelerate Goal-Conditioned RL. In this framework, the teacher module dynamically prioritizes goals with the highest temporal variance in the policy's confidence score, parameterized by the state-action value (Q) function. The teacher provides an adaptive and focused learning signal by targeting these high-uncertainty goals, fostering continual and efficient progress. We establish a theoretical connection between the temporal variance of Q-values and the evolution of the policy, providing insights into the method's underlying principles. Our approach is algorithm-agnostic and integrates seamlessly with existing RL frameworks. We demonstrate this through evaluation across 11 diverse robotic manipulation and maze navigation tasks. The results show consistent and notable improvements over state-of-the-art curriculum learning and goal-selection methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。