用单一策略学习多目标权衡,实现稠密帕累托前沿覆盖。
A Single Deep Preference-Conditioned Policy for Learning Pareto Coverage Sets

- 基于平滑切比雪夫标量化,证明偏好与最优回报一一对应且连续。
- 算法收敛速度达O(1/k),在8个任务上超主流基线的超体积排名。
- 适用于需要连续权衡探索的多目标强化学习场景。
偏好条件化的多目标强化学习旨在学习一个能捕捉偏好间权衡的单一策略,但在非线性标量化下,偏好与解之间的唯一性和连续性仍不明确。本文在表格型多目标马尔可夫决策过程(MDPs)中,采用光滑切比雪夫标量化作为单调效用函数,证明在偏好集满足弱内点条件下,每个偏好对应唯一的帕累托最优回报向量,且该向量关于偏好呈Lipschitz连续,为偏好扫描实现稠密帕累托前沿覆盖提供了理论基础。为计算这些目标,我们基于占用测度建模问题,提出凹镜面下降策略迭代(CMDPI),实现O(1/k)的期望目标次优率。进一步证明每次更新等价于求解以先前策略为参考的KL正则化MDP,获得策略迭代解释和跨偏好的有限迭代策略连续性。我们将更新实例化为深度演员-评论家算法,保留前策略正则化。在八个MO-Gymnasium任务上,该方法在平均超体积排名上优于近期基线,并表现出强预期效用性能;连续控制实验显示其优势超越离散动作设定。
原文摘要 · Abstract (English)
Preference-conditioned multi-objective reinforcement learning aims to learn a single policy that captures trade-offs across preferences, but under nonlinear scalarization the uniqueness and continuity of the preference-to-solution correspondence remain unclear. We study this problem in tabular multi-objective Markov decision processes (MDPs) using smooth Tchebycheff scalarization as a monotone utility. Under mild interior conditions on the preference set, we prove that each preference induces a unique Pareto-optimal return vector and that this vector depends Lipschitz-continuously on the preference, providing a principled foundation for preference sweeping toward dense Pareto-front coverage. To compute these targets, we formulate the problem over occupancy measures and derive Concave Mirror Descent Policy Iteration (CMDPI), which achieves an $O(1/k)$ objective-suboptimality rate. We further show that each update is equivalent to solving a Kullback-Leibler-regularized MDP with the previous policy as reference, yielding a policy-iteration interpretation and finite-iterate policy continuity across preferences. We instantiate the update as a deep actor-critic algorithm preserving previous-policy regularization. On eight MO-Gymnasium tasks, it achieves the best average hypervolume rank among recent baselines and strong expected-utility performance. Continuous-control experiments indicate gains beyond the discrete-action setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。