通过融合金融、数学和日语专家模型,打造专业金融大模型。
Merging Continual Pretraining Models for Domain-Specialized LLMs: A Case Study in Finance
- 分三阶段评估知识恢复、互补性和跨域涌现能力,验证融合有效性。
- 融合专家模型后性能提升,部分组合出现跨领域新技能。
- TIES方法更稳定,任务敏感度低,适合实际部署。
尽管大语言模型在通用任务上表现优异,但在金融等专业领域仍面临挑战,需具备领域知识、数学推理和多语言处理等多种能力。合并领域专用的持续预训练(CPT)‘专家’模型,为避免昂贵且不稳定的多技能训练提供了可行替代方案。然而,与成熟的监督微调(SFT)模型融合不同,CPT模型融合仍鲜有研究。本文以金融、数学和日语三个专家模型为基础,构建金融领域大模型。提出包含知识恢复、互补性与涌现性在内的三阶段评估框架,并在涵盖18个任务、来自8个公开数据集的综合性金融基准上,评估三种融合方法(任务算术、TIES、DARE-TIES)。结果表明:将专家模型与基础模型融合可恢复持续预训练中丢失的一般知识;多专家融合能提升性能并产生跨领域涌现能力。其中,任务算术表现良好但对超参数敏感,而TIES更具鲁棒性。研究发现模型相似性虽与融合成功率相关,但涌现能力受更复杂因素影响。本工作首次系统分析了CPT模型融合,建立原则性框架,为利用现有模型资产构建多技能大模型提供清晰指导。
原文摘要 · Abstract (English)
While LLMs excel at general tasks, they struggle in specialized domains like finance, requiring diverse skills in domain knowledge, mathematical reasoning, and multilingual processing. Merging domain-specific Continual Pre-training (CPT) "experts" offers a practical alternative to costly and unstable multi-skill training. However, unlike established Supervised Fine-Tuning (SFT) model-based merging, CPT model merging remains largely unexplored. We address this gap by creating financial LLMs from experts in finance, math, and Japanese. We propose a three-stage evaluation focusing on knowledge recovery, complementarity, and emergence, and assess three merging methods (Task Arithmetic, TIES, and DARE-TIES) on a comprehensive financial benchmark curated from 18 tasks across 8 established datasets. Results show that merging an expert with its base model recovers general knowledge lost during CPT, while merging experts improves performance and can yield emergent cross-domain skills. Among the methods, Task Arithmetic performs strongly but is hyperparameter-sensitive, whereas TIES is more robust. Our findings also suggest that while model similarity correlates with merging success, emergent skills depend on more complex factors. This work presents the first foundational analysis of CPT model merging, establishing a principled framework and providing clear guidance for building multi-skill LLMs from existing assets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。