用模型融合保留通用能力,高效适配金融大模型。
SPEAR-MM: Selective Parameter Evaluation and Restoration via Model Merging for Efficient Financial LLM Adaptation
- 通过后验分析选择性冻结或恢复层,避免遗忘通用推理能力。
- 金融任务上保留91.2%通用能力,比传统方法高21.5个百分点。
- 计算成本降低90%,适合资源受限的金融机构部署。
将大语言模型适配金融领域时,常因灾难性遗忘而丧失客户交互与复杂金融分析所需的通用推理能力。本文提出基于模型融合的可选参数评估与恢复框架SPEAR-MM,通过后验分析估算各层对基准测试的影响,再利用球面插值合并技术选择性冻结或恢复Transformer层。在LLaMA-3.1-8B上应用于金融任务时,SPEAR-MM实现91.2%的通用能力保留率,显著优于标准持续预训练的69.7%;同时保持94%的领域适应收益。该方法提供可解释的权衡控制,并将计算成本降低90%,对资源受限的金融机构极具价值。
原文摘要 · Abstract (English)
Large language models (LLMs) adapted to financial domains often suffer from catastrophic forgetting of general reasoning capabilities essential for customer interactions and complex financial analysis. We introduce Selective Parameter Evaluation and Restoration via Model Merging (SPEAR-MM), a practical framework that preserves critical capabilities while enabling domain adaptation. Our method approximates layer-wise impact on external benchmarks through post-hoc analysis, then selectively freezes or restores transformer layers via spherical interpolation merging. Applied to LLaMA-3.1-8B for financial tasks, SPEAR-MM achieves 91.2% retention of general capabilities versus 69.7% for standard continual pretraining, while maintaining 94% of domain adaptation gains. The approach provides interpretable trade-off control and reduces computational costs by 90% crucial for resource-constrained financial institutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。