arXiv:2605.20619cs.LGmath.OC2026-05

让多目标优化均匀覆盖最优解前沿,提升解的多样性。

SURF: Steering the Scalarization Weight to Uniformly Traverse the Pareto Front

论文配图:SURF: Steering the Scalarization Weight to Uniformly Traverse the Pareto Front
图 1 · 摘自论文原文
  • 通过几何分析设计权重采样策略,实现均匀遍历帕累托前沿。
  • 在多目标强化学习与大模型对齐中,解分布均匀性优于基线方法。
  • 适用于结构化与通用问题,理论保证收敛速度与采样质量。

标量法因简洁性和可扩展性被广泛用于多目标优化。许多应用需要生成反映多样化用户偏好的解,理想情况是均匀覆盖帕累托前沿(PF)。然而,均匀采样标量权重通常导致PF覆盖非均匀。本文通过标量路径的几何分析揭示这一失配:当权重变化时,对应解沿PF以非均匀速度移动,其弧长累积分布函数(CDF)决定了采样偏差。反演该CDF可得实现均匀覆盖的权重选择准则。基于此,提出SURF(Sampling Uniformly along the PaReto Front)。对于结构化问题(如双目标老虎机),推导出该CDF及其对应的权重采样规则的闭式表达;对一般问题,采用交替进行CDF重构与权重采样的迭代方法。理论上,在可证明条件下,SURF线性收敛至不可避免的有限采样下界。实验表明,在老虎机、多目标Gymnasium及多目标LLM对齐任务中,SURF能更高效地实现比基线方法更均匀的帕累托前沿覆盖。

原文摘要 · Abstract (English)

Scalarization is widely used in multi-objective optimization owing to its simplicity and scalability. In many applications, the goal is to generate solutions that represent diverse user preferences, ideally with uniform coverage of the Pareto front (PF). However, uniformly sampling scalarization weights usually induces non-uniform coverage of the PF. We explain this mismatch through a geometric analysis of the scalarization path. As the scalarization weight varies, the corresponding solutions trace the PF with a generally non-uniform traversal speed. This speed induces an arc-length cumulative distribution function (CDF); inverting this CDF map yields a principled rule for selecting weights that produce uniform PF coverage. Building on this insight, we propose SURF (Sampling Uniformly along the PaReto Front). For structured problems, including bi-objective bandits, we derive closed-form expressions for this CDF map and the resulting PF-aware weight sampling rule. For general problems, SURF alternates between CDF reconstruction and weight sampling. Theoretically, we show that under provable conditions, SURF converges linearly to an unavoidable finite-sampling floor. Empirically, experiments on bandits, multi-objective-gymnasium, and multi-objective LLM alignment demonstrate that SURF efficiently achieves more uniform PF coverage than baselines.

多目标优化帕累托前沿强化学习采样策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。