提出无需目标网络的分布值估计方法,显著提升质量多样性算法的样本效率。
Distributional Value Estimation Without Target Networks for Robust Quality-Diversity

- 采用无目标网络的分布值估计,实现高更新率训练
- 在Brax环境中用少10倍样本达到相近覆盖与性能
- 适合追求高效进化强化学习的研究者
质量多样性(QD)算法擅长发现多样化的技能集合,但受限于低样本效率,通常需数千万环境步才能解决复杂运动任务。近期强化学习进展表明,高更新-数据比(UTD)可加速演员-评论家学习,但标准高UTD方法依赖目标网络稳定训练,造成显著计算瓶颈,难以应用于对样本效率和快速种群适应性要求高的质量多样性任务。本文提出QDHUAC,一种样本高效、无目标网络且具备分布值估计的质量多样性强化学习算法,提供密集且低方差梯度信号,支持高UTD训练下的支配性新颖性搜索。实验表明,该方法可在高UTD下实现稳定训练,在高维Brax环境中以约十分之一的样本量达成与基线相当的覆盖率与适应度。结果表明,将无目标分布评论家与基于支配的选择机制结合,是下一代高效进化强化学习的关键。
原文摘要 · Abstract (English)
Quality-Diversity (QD) algorithms excel at discovering diverse repertoires of skills, but are hindered by poor sample efficiency and often require tens of millions of environment steps to solve complex locomotion tasks. Recent advances in Reinforcement Learning (RL) have shown that high Update-to-Data (UTD) ratios accelerate Actor-Critic learning. While effective, standard high-UTD algorithms typically utilise target networks to stabilise training. This requirement introduces a significant computational bottleneck, rendering them impractical for resource-intensive Quality-Diversity (QD) tasks where sample efficiency and rapid population adaptation are critical. In this paper, we introduce QDHUAC, a sample-efficient, target-free and distributional QD-RL algorithm that provides dense and low-variance gradient signals, which enables high-UTD training for Dominated Novelty Search whilst requiring an order of magnitude fewer environment steps. We demonstrate that our method enables stable training at high UTD ratios, achieving competitive coverage and fitness on high-dimensional Brax environments with an order of magnitude fewer samples than baselines. Our results suggest that combining target-free distributional critics with dominance-based selection is a key enabler for the next generation of sample-efficient evolutionary RL algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。