arXiv:2506.05980cs.LGcs.AI2025-06

提出AMPED方法,让技能学习兼顾探索与多样性。

AMPED: Adaptive Multi-objective Projection for balancing Exploration and skill Diversification

  • 用梯度修正平衡探索与多样性梯度
  • 在多个基准上超越现有基线表现
  • 适合需要泛化技能的学习任务

基于技能的强化学习(SBRL)通过预训练技能条件策略,在稀疏奖励环境中实现快速适应。有效的技能学习需同时最大化探索和技能多样性,但现有方法难以兼顾这两个相互冲突的目标。本文提出自适应多目标投影方法AMPED,显式解决该问题:预训练阶段采用梯度手术投影平衡探索与多样性梯度,微调阶段使用技能选择器根据下游任务挑选合适技能。实验表明,AMPED在多个基准上优于现有SBRL基线。消融研究验证了各组件贡献,且理论与实证证明,使用贪心技能选择器时,更高的技能多样性可降低微调样本复杂度。结果强调了显式协调探索与多样性的关键作用,展示了AMPED在实现鲁棒、泛化性强的技能学习中的有效性。

原文摘要 · Abstract (English)

Skill-based reinforcement learning (SBRL) enables rapid adaptation in environments with sparse rewards by pretraining a skill-conditioned policy. Effective skill learning requires jointly maximizing both exploration and skill diversity. However, existing methods often face challenges in simultaneously optimizing for these two conflicting objectives. In this work, we propose a new method, Adaptive Multi-objective Projection for balancing Exploration and skill Diversification (AMPED), which explicitly addresses both: during pre-training, a gradient-surgery projection balances the exploration and diversity gradients, and during fine-tuning, a skill selector exploits the learned diversity by choosing skills suited to downstream tasks. Our approach achieves performance that surpasses SBRL baselines across various benchmarks. Through an extensive ablation study, we identify the role of each component and demonstrate that each element in AMPED is contributing to performance. We further provide theoretical and empirical evidence that, with a greedy skill selector, greater skill diversity reduces fine-tuning sample complexity. These results highlight the importance of explicitly harmonizing exploration and diversity and demonstrate the effectiveness of AMPED in enabling robust and generalizable skill learning. Project Page: https://geonwoo.me/amped/

强化学习技能学习多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。