用变分合并方法更精准估计多任务微调的帕累托前沿。
Variational Model Merging for Pareto Front Estimation in Multitask Finetuning
- 基于贝叶斯框架设计变分模型合并,灵活后验提升合并质量。
- 非球面高斯后验可显著改善帕累托前沿估计效果。
- 适用于需权衡多任务性能的研究者,尤其关注模型融合优化。
帕累托前沿有助于发现多任务微调中的最优任务混合策略,但其计算成本较高。现有工作通过模型合并构建低成本代理模型来估计帕累托前沿,但尚未有研究专门设计新合并方法以直接且理论上提升前沿质量。本文提出一种新的贝叶斯方法——变分模型合并(Variational Model Merging),将已有合并方法视为使用高斯后验时的特例,而通过采用非高斯后验可推导出新型合并策略。理论核心结果表明,更灵活的后验必然带来更好的帕累托前沿估计。例如,使用全高斯后验的合并结果优于各向同性高斯后验。我们在视觉与语言变压器模型上进行了广泛实验验证,结果表明更复杂的高斯族始终能获得更优或相当的帕累托前沿。本工作是少数将贝叶斯思想用于改进帕累托分析的案例。
原文摘要 · Abstract (English)
Pareto fronts are useful to find good task-mixing strategies for multitask finetuning, but they are also costly to compute. To reduce costs, recent works have used existing model merging methods to help train cheap surrogate models to estimate the Pareto fronts. However, no work has yet considered designing new model-merging methods to directly, and provably, improve the quality of Pareto fronts. Here, we fill this gap by proposing a new Bayesian approach called Variational Model Merging. In this approach, existing model-merging methods are obtained as special cases of "posterior-merging" when Gaussian posteriors are used and new model-merging strategies can be derived by using non-Gaussian posteriors. Our main theoretical result is to show that more flexible posteriors necessarily yield better estimates of Pareto fronts. For instance, a Pareto front estimate obtained by merging full-Gaussian posteriors is expected to be better than that obtained by using isotropic Gaussian posteriors. We validate the theory through extensive empirical results on vision and language transformers where better Gaussian families consistently yields better or comparable Pareto fronts. Our work is a rare instance where Bayesian ideas are used to improve Pareto analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。