arXiv:2604.19394cs.CL2026-04ACL被引 1

用持续预训练让小模型变强,医疗任务表现接近大模型。

Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?

论文配图:Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
图 1 · 摘自论文原文
  • 通过持续预训练和模型融合,提升小模型在医学领域的表现。
  • 70亿参数模型经改造后,在德语医疗数据上性能大幅提升,胜率超大模型3.5倍。
  • 适合资源有限但需高精度医疗问答的场景,如德国医疗机构。

本文通过持续预训练与模型融合,缩小了小型专用模型与大型通用模型在医学领域的性能差距。为解决非英语医学数据稀缺问题,研究者从FineWeb2构建了一个高质量德语医学语料库(FineMed-de),用于持续预训练并融合三款知名大模型(参数量70亿至240亿)。实验表明,专业化显著提升了70亿参数模型在德语医疗基准上的表现。基于Qwen2.5的模型对战分析显示,其胜率相较更大的Mistral-Small-24B-Instruct模型提升约3.5倍。结果表明,经过领域适配的70亿参数模型可成为复杂医疗指令遵循任务中高效且具竞争力的解决方案。尽管模型融合恢复了指令跟随能力,但失败模式分析揭示存在语言混杂与表达冗余等固有权衡,提示未来需更精准微调。该研究提供了符合实际需求、可合规部署的专用大模型开发方法,为德语医疗场景落地奠定基础。

原文摘要 · Abstract (English)

This paper narrows the performance gap between small, specialized models and significantly larger general-purpose models through domain adaptation via continual pre-training and merging. We address the scarcity of specialized non-English data by constructing a high-quality German medical corpus (FineMed-de) from FineWeb2. This corpus is used to continually pre-train and merge three well-known LLMs (ranging from $7B$ to $24B$ parameters), creating the DeFineMed model family. A comprehensive evaluation confirms that specialization dramatically enhances $7B$ model performance on German medical benchmarks. Furthermore, the pairwise win-rate analysis of the Qwen2.5-based models demonstrates an approximately $3.5$-fold increase in the win-rate against the much larger Mistral-Small-24B-Instruct through domain adaptation. This evidence positions specialized $7B$ models as a competitive, resource-efficient solution for complex medical instruction-following tasks. While model merging successfully restores instruction-following abilities, a subsequent failure mode analysis reveals inherent trade-offs, including the introduction of language mixing and increased verbosity, highlighting the need for more targeted fine-tuning in future work. This research provides a robust, compliant methodology for developing specialized LLMs, serving as the foundation for practical use in German-speaking healthcare contexts.

医学大模型模型融合持续预训练德语NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。