arXiv:2603.14621eess.IVcs.CV2026-03

多中心新冠肺部CT分类,用混合模型+校准提升诊断鲁棒性

A Heterogeneous Ensemble for Multi-Center COVID-19 Classification from Chest CT Scans

  • 九种模型融合三种不同思路,跨中心适应性强
  • 验证损失比从35倍降至3倍内,有效缓解过拟合
  • 适合医疗影像跨机构部署,对数据差异敏感场景

新冠疫情暴露了诊断流程的局限:RT-PCR检测耗时长、假阴性率高,而基于CT的筛查虽可快速辅助诊断,却依赖专业放射科医生解读。在多个医院中心部署自动化CT分析时,由于扫描仪硬件、采集协议和患者群体差异导致显著领域偏移,单模型性能下降。为此,我们提出一个包含九个模型的异构集成方法,涵盖三种推理范式:(1) 自监督DINOv2视觉变换器结合切片级sigmoid聚合,(2) RadImageNet预训练DenseNet-121采用切片级sigmoid平均,(3) 七种基于EfficientNet-B3、ConvNeXt-Tiny和EfficientNetV2-S骨干网络的门控注意力多实例学习模型,使用扫描级softmax分类。通过随机种子变化与随机权重平均增强集成多样性。利用焦点损失、嵌入层混补和领域感知增强,将验证集与训练集损失比从35倍降低至不足3倍,有效缓解严重过拟合。模型输出通过加权概率平均融合,并采用各源独立阈值优化进行校准。最终集成模型在四个医院中心上实现平均宏F1为0.9280,优于最佳单模型(F1=0.8969)的+0.031,表明异构架构结合源感知校准对于多中心医学图像分类至关重要。

原文摘要 · Abstract (English)

The COVID-19 pandemic exposed critical limitations in diagnostic workflows: RT-PCR tests suffer from slow turnaround times and high false-negative rates, while CT-based screening offers faster complementary diagnosis but requires expert radiological interpretation. Deploying automated CT analysis across multiple hospital centres introduces further challenges, as differences in scanner hardware, acquisition protocols, and patient populations cause substantial domain shift that degrades single-model performance. To address these challenges, we present a heterogeneous ensemble of nine models spanning three inference paradigms: (1) a self-supervised DINOv2 Vision Transformer with slice-level sigmoid aggregation, (2) a RadImageNet-pretrained DenseNet-121 with slice-level sigmoid averaging, and (3) seven Gated Attention Multiple Instance Learning models using EfficientNet-B3, ConvNeXt-Tiny, and EfficientNetV2-S backbones with scan-level softmax classification. Ensemble diversity is further enhanced through random-seed variation and Stochastic Weight Averaging. We address severe overfitting, reducing the validation-to-training loss ratio from 35x to less than 3x, through a combination of Focal Loss, embedding-level Mixup, and domain-aware augmentation. Model outputs are fused via score-weighted probability averaging and calibrated with per-source threshold optimization. The final ensemble achieves an average macro F1 of 0.9280 across four hospital centres, outperforming the best single model (F1=0.8969) by +0.031, demonstrating that heterogeneous architectures combined with source-aware calibration are essential for robust multi-site medical image classification.

多中心分类医学影像模型集成新冠筛查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。