用偏好引导优化多目标模型,高效找到更优解
Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference Guidance
- 将多目标约束转为单目标,通过可计算梯度的函数实现
- 在合成与真实任务中均有效提升偏好引导下的模型性能
- 适合需要兼顾多个目标且有明确偏好的实际应用
在用户指定偏好下的多目标学习广泛存在于现实问题中,如公平性要求下的多语言语音识别。本文将此类问题建模为半向量双层优化问题,目标是优化预定义的偏好函数,同时要求模型参数处于弱帕累托最优状态。为此,我们通过一个具有易计算梯度的业绩函数,将多目标约束转化为单目标约束,并采用基于惩罚项的双层优化重构方法。理论上分析了业绩函数性质及惩罚重构解与约束解之间的关系。进一步提出求解重构后单层问题的算法,并给出收敛性保证。在多种合成与真实场景中测试表明,该方法能有效找到偏好引导下的多目标最优解。
原文摘要 · Abstract (English)
Multi-objective learning under user-specified preference is common in real-world problems such as multi-lingual speech recognition under fairness. In this work, we frame such a problem as a semivectorial bilevel optimization problem, whose goal is to optimize a pre-defined preference function, subject to the constraint that the model parameters are weakly Pareto optimal. To solve this problem, we convert the multi-objective constraints to a single-objective constraint through a merit function with an easy-to-evaluate gradient, and then, we use a penalty-based reformulation of the bilevel optimization problem. We theoretically establish the properties of the merit function, and the relations of solutions for the penalty reformulation and the constrained formulation. Then we propose algorithms to solve the reformulated single-level problem, and establish its convergence guarantees. We test the method on various synthetic and real-world problems. The results demonstrate the effectiveness of the proposed method in finding preference-guided optimal solutions to the multi-objective problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。