arXiv:2604.01538cs.CLcs.AI2026-04

用模型合并缓解大模型医嘱遗忘,提升临床应用能力。

Countering Catastrophic Forgetting of Large Language Models for Better Instruction Following via Weight-Space Model Merging

  • 通过权重空间插值合并通用指令模型与医疗基础模型
  • 在64样本下性能媲美256样本微调,保留医嘱遵循能力
  • 适合资源有限的医疗环境快速部署开源大模型

大型语言模型在医疗临床文档中被用于减轻医生负担,但研究发现其在使用特定医疗数据集微调时会严重丧失指令遵循能力,成为通用大模型应用于临床的关键挑战。本研究提出一种模型合并框架,通过插值方法将临床基础模型(GatorTronLlama)与通用指令模型(Llama-3.1-8B-Instruct)融合,有效应对该遗忘问题。在多个医学基准和五项临床生成任务(如放射科报告、出院摘要)上的评估表明,合并模型能显著缓解灾难性遗忘,保持临床领域专长并保留指令遵循能力。此外,该策略在极低监督条件下(如64样本)即达到与全量微调基线(256样本)相当的性能,展现出高训练效率。因此,权重空间合并为开源大模型在临床场景中的可扩展适应提供了可行方案,助力资源受限的医疗环境实现更广泛部署。

原文摘要 · Abstract (English)

Large language models have been adopted in the medical domain for clinical documentation to reduce clinician burden. However, studies have reported that LLMs often "forget" a significant amount of instruction-following ability when fine-tuned using a task-specific medical dataset, a critical challenge in adopting general-purpose LLMs for clinical applications. This study presents a model merging framework to efficiently adapt general-purpose LLMs to the medical domain by countering this forgetting issue. By merging a clinical foundation model (GatorTronLlama) with a general instruct model (Llama-3.1-8B-Instruct) via interpolation-based merge methods, we seek to derive a domain-adapted model with strong performance on clinical tasks while retaining instruction-following ability. Comprehensive evaluation across medical benchmarks and five clinical generation tasks (e.g., radiology and discharge summarization) shows that merged models can effectively mitigate catastrophic forgetting, preserve clinical domain expertise, and retain instruction-following ability. In addition, our model merging strategies demonstrate training efficiency, achieving performance on par with fully fine-tuned baselines under severely constrained supervision (e.g., 64-shot vs. 256-shot). Consequently, weight-space merging constitutes a highly scalable solution for adapting open-source LLMs to clinical applications, facilitating broader deployment in resource-constrained healthcare environments.

大模型微调医疗AI模型合并指令遵循

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。