通过结构化分支联合学习策略,生成高质量且多样化的智能体
Structure-Conditioned Actor-Critic Branches for Quality-Diversity Reinforcement Learning

- 用结构掩码限定策略学习空间,每条分支独立训练价值函数
- 在MuJoCo上构建的策略库兼具高奖励与行为多样性
- 适合需要动态调整行为目标的连续控制场景
质量-多样性强化学习(QD-RL)旨在构建包含高性能与行为多样性的策略集合。现有方法多在策略执行后进行多样性拓展,或利用价值信息提升策略质量,但对生成候选策略的学习分支研究不足。本文提出SV-QD-RL,一种结构-价值耦合框架,将每个候选策略表示为结构条件化的演员-评论家分支。每个分支包含演员、结构掩码、分支特异性评论家、回放缓冲区及行为、回报、稀疏性与价值分布等评估属性。结构掩码定义演员的学习子空间,分支特异性评论家与回放缓冲区塑造其价值学习轨迹。一个分支感知的QD归档根据行为质量、结构痕迹与价值分布信息评估并保留分支。在MuJoCo连续控制任务上的实验表明,SV-QD-RL构建的策略库具有优异的归档质量与有用的行为多样性。消融与诊断分析进一步显示,结构条件化、评论家差异化与记忆一致性优化对行为专业化有互补贡献。时序感知的策略库评估表明,学习到的归档可在行为需求变化时提供可选策略。结果表明,将演员结构与分支特异性价值学习耦合,是生成多样化QD-RL策略集合的有效机制。
原文摘要 · Abstract (English)
Quality-diversity reinforcement learning (QD-RL) aims to construct policy repertoires that contain both high-performing and behaviorally diverse policies. Existing QD-RL methods mainly diversify policy instances after rollout evaluation or use learned value information to improve policy quality and behavior targeting, while the learning branches that generate candidate policies remain less explored. This paper proposes SV-QD-RL, a structure-value coupled framework that represents each candidate as a structure-conditioned actor-critic branch. Each branch contains an actor, a structural mask, a branch-specific critic, a replay state, and evaluation attributes including behavior, return, sparsity, and value profile. The structural mask defines the actor subspace in which the branch learns, while the branch-specific critic and replay state shape its value-learning trajectory. A branch-aware QD archive then evaluates and retains branches according to behavioral quality, structural footprint, and value-profile information. Experiments on MuJoCo continuous-control tasks show that SV-QD-RL constructs policy repertoires with strong archive quality and behaviorally useful diversity. Ablation and diagnostic analyses further indicate that structural conditioning, critic differentiation, and memory-consistent refinement make complementary contributions to behavioral specialization. Schedule-aware repertoire evaluation shows that the learned archive provides selectable policy alternatives under changing behavior-level requirements. These results suggest that coupling actor structure with branch-specific value learning is an effective mechanism for generating diverse QD-RL policy repertoires.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。