自动合并多个大模型,用少算力找到更优组合。
Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging
- 用多保真度搜索自动优化合并策略,省去人工调参。
- 在少量计算(<500步)内提升单一任务和多任务性能。
- 适合想高效融合模型但缺乏资源的研究者或工程师。
推理能力是大型语言模型的关键挑战,但开发需大量专有数据和算力。模型合并提供了一种无需重训练即可融合多个模型的可行方案。然而,现有方法依赖人工设计的合并超参数,限制了组合探索并需大量人力。本文提出自动化模型合并框架,通过多保真度近似实现细粒度策略探索,降低计算成本。支持单目标与多目标优化,引入两种新搜索空间:层级融合(LFS)和深度整合(DIS)。在多个基准测试中,该方法自主发现:1)进一步提升单一任务性能,包括模型已微调过的任务;2)优化跨任务的多目标性能前沿。有效合并可在少于500次搜索步骤内完成。
原文摘要 · Abstract (English)
Reasoning capabilities represent a critical frontier for large language models (LLMs), but developing them requires extensive proprietary datasets and computational resources. One way to efficiently supplement capabilities with is by model merging, which offers a promising alternative by combining multiple models without retraining. However, current merging approaches rely on manually-designed strategies for merging hyperparameters, limiting the exploration of potential model combinations and requiring significant human effort. We propose an Automated Model Merging Framework that enables fine-grained exploration of merging strategies while reducing costs through multi-fidelity approximations. We support both single and multi-objective optimization and introduce two novel search spaces: layerwise fusion (LFS) and depth-wise integration (DIS). Evaluating across a number of benchmarks, we find that the search autonomously finds 1) Merges that further boost single-objective performance, even on tasks the model has already been finetuned on, and 2) Merges that optimize multi-objective frontiers across tasks. Effective merges are found with limited compute, e.g. within less than 500 search steps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。