arXiv:2409.03444cs.CLcond-mat.mtrl-sci2024-09被引 167

通过模型合并实现领域大模型能力跃升,小模型也能变聪明。

Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities

论文配图:Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities
图 1 · 摘自论文原文
  • 用多种微调策略训练模型,再合并多个模型提升性能
  • 合并后模型能力超越单个模型,出现新功能
  • 小模型合并后仍可表现优异,适合资源有限场景

大型语言模型(LLMs)在材料科学与工程等领域的应用依赖于针对特定领域的微调策略。本文研究了持续预训练(CPT)、监督微调(SFT)及基于偏好的优化方法(如直接偏好优化DPO、奇数比偏好优化ORPO)对微调后模型性能的影响。结果表明,多模型合并不仅能提升性能,还能催生单个模型不具备的新能力。实验使用Llama 3.1 8B和Mistral 7B模型,均观察到类似现象;进一步测试17亿参数的小模型,发现其合并后未表现出涌现能力,说明模型规模可能是关键因素。在开放对话中,最小模型在推理深度、创意、清晰度和定量精度等维度均获得高智能评分。此外,基于生物材料设计理念生成图像生成提示,成功设计出新型微观结构、建筑概念与城市规划方案。

原文摘要 · Abstract (English)

The advancement of Large Language Models (LLMs) for domain applications in fields such as materials science and engineering depends on the development of fine-tuning strategies that adapt models for specialized, technical capabilities. In this work, we explore the effects of Continued Pretraining (CPT), Supervised Fine-Tuning (SFT), and various preference-based optimization approaches, including Direct Preference Optimization (DPO) and Odds Ratio Preference Optimization (ORPO), on fine-tuned LLM performance. Our analysis shows how these strategies influence model outcomes and reveals that the merging of multiple fine-tuned models can lead to the emergence of capabilities that surpass the individual contributions of the parent models. We find that model merging leads to new functionalities that neither parent model could achieve alone, leading to improved performance in domain-specific assessments. Experiments with different model architectures are presented, including Llama 3.1 8B and Mistral 7B models, where similar behaviors are observed. Exploring whether the results hold also for much smaller models, we use a tiny LLM with 1.7 billion parameters and show that very small LLMs do not necessarily feature emergent capabilities under model merging, suggesting that model scaling may be a key component. In open-ended yet consistent chat conversations between a human and AI models, our assessment reveals detailed insights into how different model variants perform and show that the smallest model achieves a high intelligence score across key criteria including reasoning depth, creativity, clarity, and quantitative precision. Other experiments include the development of image generation prompts based on disparate biological material design concepts, to create new microstructures, architectural concepts, and urban design based on biological materials-inspired construction principles.

大模型微调模型合并领域适应小模型智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。