0.1B参数的开源医学模型,小规模下展现独特能力分化。
MedLLM: An Open Medical Language Model at the Sub-Billion Scale

- 三阶段训练:通用预训练+医学语料微调+偏好对齐
- 小模型在问答任务中接近7B模型表现,知识召回仍具优势
- 适合研究小模型医学能力边界或资源受限场景部署
开放医学语言模型普遍采用70亿参数以上规模,子十亿级别仍缺乏系统研究。本文提出MedLLM,一个0.1B参数的开源医学语言模型,通过全开源三阶段流程训练:基于课程序列长度调度的通用预训练、在MedFineWeb上的领域微调(从通用网页数据中筛选与医学问答相似的文本),以及结合SFT与直接偏好优化(DPO)的偏好对齐。在多个医学基准测试中,MedLLM展现出仅在子十亿规模下可见的现象:医学能力在压缩下并非均匀退化,而是依任务类型分化。在上下文依赖型问答任务上,其性能仅比经过医学适配的7B模型低2.9个百分点,优于所有指令微调及通用7B基线;在知识回忆型问答任务中,于临床病例题集MedQA上接近任务下限,但在MedMCQA上显著超越所有7B及子7B基线,表明此时瓶颈是模型容量而非适配性。该差异在7B模型中被掩盖,仅在容量受限时显现。
原文摘要 · Abstract (English)
Open medical language models have converged on a single scale: every widely used system runs at 7B parameters or more, leaving the sub-billion regime uncharacterized. We present MedLLM, an open 0.1B-parameter medical language model trained through a fully open three-phase pipeline: general pretraining with curriculum sequence-length scheduling, domain fine-tuning on MedFineWeb, a reference-guided medical corpus we release that is selected from general web data by embedding similarity to medical question-answering (QA) data, and preference-aligned fine-tuning combining SFT with direct preference optimization (DPO). Across medical benchmarks, MedLLM shows a pattern visible only at sub-billion scale: medical competence does not degrade uniformly under compression but splits by task type. On context-grounded QA it comes within $2.9$pp of a medically adapted 7B model and surpasses the instruction-tuned and general-purpose 7B baselines; on knowledge-recall QA it stays near the task floor on clinical-vignette MedQA yet significantly exceeds every 7B and sub-7B baseline on MedMCQA, indicating that where recall fails the constraint is model capacity rather than adaptation. This dissociation is masked at 7B, where both capabilities are present, and surfaces only when capacity is scarce.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。