无需判别器,用新机制提升离线强化学习多样性与稳定性
Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints
- 用后续特征计算范德华力目标,避免依赖不稳定判别器
- 在两个模拟任务中实现高多样性且高性能行为,匹配专家状态分布
- 支持零样本技能召回,无需预先设定技能数量
在模仿约束下的离线多样性最大化可将示范数据转化为一组不同的行为策略,提升对分布偏移的鲁棒性,且无需额外环境交互。然而,现有方法常依赖需要训练技能判别器的互信息目标,在交替拉格朗日优化引发的非平稳奖励下容易不稳定。本文提出 Dual-Force,一种离线算法:(i) 利用从后续特征计算的非监督范德华力(VdW)目标最大化多样性,无需技能判别器;(ii) 通过预训练的功能奖励编码(FRE)条件化价值函数与策略,稳定非平稳内在奖励下的训练过程。FRE 编码还支持通过关联隐表示实现零样本技能召回,无需预设技能数量。在两个 Solo12 模拟基准(运动与障碍导航)上,Dual-Force 在匹配目标专家状态占用率的同时,恢复多样且高性能的行为,并在对抗性障碍变化中表现更鲁棒。
原文摘要 · Abstract (English)
Offline diversity maximization under imitation constraints can transform demonstration data into a set of distinct behavioral policies, improving robustness to distribution shift without additional environment interaction. In practice, however, existing offline approaches often rely on mutual-information objectives that require training a skill discriminator and can become unstable under the non-stationary rewards induced by alternating Lagrangian optimization. We introduce Dual-Force, an offline algorithm that (i) maximizes diversity using an off-policy estimator of a Van der Waals (VdW) force objective computed from successor features, eliminating the skill discriminator, and (ii) stabilizes training under non-stationary intrinsic rewards by conditioning the value function and policy on a pre-trained Functional Reward Encoding (FRE). The FRE code also enables zero-shot recall of every encountered skill via its associated latent representation, removing the need to pre-specify a fixed number of skills. On two Solo12 simulation benchmarks (locomotion and obstacle navigation), Dual-Force recovers diverse high-performing behaviors while matching a target expert state occupancy and improves robustness in adversarial obstacle variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。