arXiv:2603.15154eess.IVcs.CV2026-03

针对多机构CT影像差异,提出分阶段专家融合模型提升新冠检测鲁棒性。

Vision-Language Model Based Multi-Expert Fusion for CT Image Classification

  • 构建肺部感知3D专家,融合原始与肺部提取图像进行分类。
  • 设计两阶段专家:逐切片表征与跨切片上下文建模,捕捉多层特征。
  • 引入源身份预测与分层投票机制,显著提升异构数据下分类性能。

由于多机构设置下的源域偏移、源域不平衡及隐藏测试源身份问题,从胸部CT中稳健检测新冠仍具挑战。本文提出一种三阶段源感知多专家框架,用于多源新冠CT分类。首先,构建结合原始CT与肺部提取CT的肺部感知3D专家,实现体积分层分类;其次,开发两个基于MedSigLIP的专家:一个用于逐切片表征与概率学习,另一个通过Transformer建模跨切片依赖关系;第三,训练源分类器以预测每个测试扫描的潜在来源身份。利用预测的源信息,对不同专家进行模型融合与投票。在涵盖四个来源的验证集上,第一阶段模型达到最高宏平均F1为0.9711,准确率0.9712,AUC为0.9791;第二阶段a和b分别取得最佳AUC 0.9864与0.9854;第三阶段源分类器准确率0.9107,F1为0.9114。结果表明,源感知专家建模与分层投票是应对异构多源条件下的有效方案。

原文摘要 · Abstract (English)

Robust detection of COVID-19 from chest CT remains challenging in multi-institutional settings due to substantial source shift, source imbalance, and hidden test-source identities. In this work, we propose a three-stage source-aware multi-expert framework for multi-source COVID-19 CT classification. First, we build a lung-aware 3D expert by combining original CT volumes and lung-extracted CT volumes for volumetric classification. Second, we develop two MedSigLIP-based experts: a slice-wise representation and probability learning module, and a Transformer-based inter-slice context modeling module for capturing cross-slice dependency. Third, we train a source classifier to predict the latent source identity of each test scan. By leveraging the predicted source information, we perform model fusion and voting based on different experts. On the validation set covering all four sources, the Stage 1 model achieves the best macro-F1 of 0.9711, ACC of 0.9712, and AUC of 0.9791. Stage~2a and Stage~2b achieve the best AUC scores of 0.9864 and 0.9854, respectively. Stage~3 source classifier reaches 0.9107 ACC and 0.9114 F1. These results demonstrate that source-aware expert modeling and hierarchical voting provide an effective solution for robust COVID-19 CT classification under heterogeneous multi-source conditions.

医学影像多源学习视觉语言模型新冠检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。