arXiv:2411.06710cs.AIcs.CL2024-11NeurIPS被引 7

用贝叶斯优化融合多个模型,提升微调效果。

Model Fusion through Bayesian Optimization in Language Model Fine-Tuning

  • 通过多目标贝叶斯优化同时优化损失和指标
  • 在多个下游任务上显著提升模型性能
  • 适合需要高效微调的NLP研究者

为下游任务微调预训练模型是一种广泛采用的技术,具有良好的适应性和可靠性。尽管概念简单,但微调涉及诸多工程决策,如超参数选择和从优化轨迹中确定检查点。为解决最佳模型选择难题,一种有效方法是模型融合,即在参数空间中组合多个模型。然而我们观察到,在预训练语言模型微调过程中,损失与指标之间的景观存在显著差异。基于此,我们提出一种新的模型融合技术,通过多目标贝叶斯优化同时优化目标指标和损失。此外,为有效选择超参数,我们在框架中引入两阶段流程,将贝叶斯优化过程集成其中。在多个下游任务上的实验表明,使用该贝叶斯优化引导的方法可带来显著性能提升。

原文摘要 · Abstract (English)

Fine-tuning pre-trained models for downstream tasks is a widely adopted technique known for its adaptability and reliability across various domains. Despite its conceptual simplicity, fine-tuning entails several troublesome engineering choices, such as selecting hyperparameters and determining checkpoints from an optimization trajectory. To tackle the difficulty of choosing the best model, one effective solution is model fusion, which combines multiple models in a parameter space. However, we observe a large discrepancy between loss and metric landscapes during the fine-tuning of pre-trained language models. Building on this observation, we introduce a novel model fusion technique that optimizes both the desired metric and loss through multi-objective Bayesian optimization. In addition, to effectively select hyperparameters, we establish a two-stage procedure by integrating Bayesian optimization processes into our framework. Experiments across various downstream tasks show considerable performance improvements using our Bayesian optimization-guided method.

模型融合贝叶斯优化微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。