提出D³PO框架,让智能体更好平衡多个冲突目标。
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization
- 分解目标优化流程,延迟引入偏好以减少信息损失
- 在多目标环境中生成更广更优的帕累托前沿
- 适合需要稳定多目标决策的强化学习应用
多目标强化学习(MORL)旨在训练能平衡冲突目标的智能体。现有基于单一偏好条件策略的方法虽具可扩展性,但实践中常因无法恢复密集帕累托前沿而失效。我们发现其根源在于两个结构性问题:过早线性化导致优势信号破坏,以及偏好空间中表示模式崩溃。为此,提出基于PPO的D³PO框架,通过分解目标学习路径、在信任区域稳定后才引入偏好(晚期加权),保留各目标学习信号。同时引入缩放多样性正则项,促进行为差异与偏好距离成比例。D³PO完全运行在标准深度MORL通用的线性加权范式内,通过减少线性加权带来的信息损失,而非依赖昂贵的非线性效用函数,表明优化瓶颈至关重要。在多个标准基准上,包括高维与多目标环境,D³PO持续生成更宽、质量更高的帕累托前沿,其超体积和期望效用均优于现有方法,仅用一个可部署策略实现。
原文摘要 · Abstract (English)
Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-conditioned policies offer a highly scalable solution, existing approaches remain brittle in practice, frequently failing to recover dense Pareto fronts. We demonstrate that this failure stems from two structural pathologies: destructive advantage cancellation caused by premature Early Scalarization (ES), and representational mode collapse across the preference space. To overcome these bottlenecks, we introduce $D^3PO$, a PPO-based framework that fundamentally reorganizes multi-objective optimization. By preserving per-objective learning signals through a decomposed pipeline and integrating preferences only after trust-region stabilization (Late-Stage Weighting), $D^3PO$ improves credit assignment under conflicting objectives. Concurrently, a scaled diversity regularizer encourages behavioral divergence proportional to preference distance. $D^3PO$ operates entirely within the efficient linear scalarization regime shared by standard deep MORL baselines. By reducing information loss caused due to linear scalarization rather than relying on expensive non-linear utility functions, it suggests that optimization bottlenecks play a significant role. Across available standard benchmarks, including high-dimensional and many-objective environments, $D^3PO$ consistently discovers broader, higher-quality Pareto fronts than prior methods, exceeding state-of-the-art hypervolume and expected utility using a single deployable policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。