用低秩专家混合提升医学影像模型对头颅CT病灶的检测能力
Specializing Foundation Models via Mixture of Low-Rank Experts for Comprehensive Head CT Analysis
- 通过多专家低秩适配器实现病灶类型自适应调整
- 在7万张CT上检测75种病变,平均准确率提升至0.917
- 仅增加0.5%参数,无需标注病灶即可自动分配适配器
大规模预训练基础模型具备强大的迁移能力,但在复杂多标签诊断任务(如全面头颅CT病灶检测)中的适应仍研究不足。标准参数高效微调方法如LoRA对所有病灶类型采用统一调整,可能限制多样病灶的表现。本文提出一种混合低秩专家(MoLRE)框架,在LoRA基础上引入多个专用低秩适配器与无监督软路由机制,实现条件化特征适配,额外参数少于0.5%,且无需显式病理标注。我们在六种先进医疗影像基础模型上进行综合评测,涵盖2D/3D架构、通用域/医学域/头颅CT特化预训练及7M至431M参数规模。基于超过7万例非增强头颅CT扫描,含75类标注病灶(包括出血、梗死、创伤、占位、结构异常和慢性改变),实验显示所有模型均获一致提升:通用与医学域模型增益最大(DINOv3-Base: +4.6%;MedGemma: +4.3%),而3D特化或超大模型提升较小(+0.2%-1.3%)。MoLRE与MedGemma结合达到最高平均检测AUC 0.917。结果凸显针对临床任务系统性评估的重要性,预训练领域、架构与模型规模间存在非直观交互效应。
原文摘要 · Abstract (English)
Foundation models pre-trained on large-scale datasets demonstrate strong transfer learning capabilities; however, their adaptation to complex multi-label diagnostic tasks-such as comprehensive head CT finding detection-remains understudied. Standard parameter-efficient fine-tuning methods such as LoRA apply uniform adaptations across pathology types, which may limit performance for diverse medical findings. We propose a Mixture of Low-Rank Experts (MoLRE) framework that extends LoRA with multiple specialized low-rank adapters and unsupervised soft routing. This approach enables conditional feature adaptation with less than 0.5% additional parameters and without explicit pathology supervision. We present a comprehensive benchmark of MoLRE across six state-of-the-art medical imaging foundation models spanning 2D and 3D architectures, general-domain, medical-domain, and head CT-specific pretraining, and model sizes ranging from 7M to 431M parameters. Using over 70,000 non-contrast head CT scans with 75 annotated findings-including hemorrhage, infarction, trauma, mass lesions, structural abnormalities, and chronic changes-our experiments demonstrate consistent performance improvements across all models. Gains vary substantially: general-purpose and medical-domain models show the largest improvements (DINOv3-Base: +4.6%; MedGemma: +4.3%), whereas 3D CT-specialized or very large models show more modest gains (+0.2-1.3%). The combination of MoLRE and MedGemma achieves the highest average detection AUC of 0.917. These findings highlight the importance of systematic benchmarking on target clinical tasks, as pretraining domain, architecture, and model scale interact in non-obvious ways.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。