基础模型在腹部创伤CT中特异性下降,主因是正常影像的多样性而非疾病稀有。
Beyond Calibration: Confounding Pathology Limits Foundation Model Specificity in Abdominal Trauma CT
- 对比基础模型与专用模型,发现特异性差异源于正常影像的异质性。
- 无腹部病变时特异性达84%-100%,但合并器官损伤时下降50-51个百分点。
- 模型越依赖标注训练,越能适应复杂负样本,适合临床部署前优化。
目的:将基础模型应用于临床需评估其在复合分布偏移下的表现,即严重类别不平衡与异质性影像并存的情况。此问题对罕见但高致死率的创伤性肠损伤尤为关键。本研究探究基础模型特异性缺陷是否与阴性类别的异质性相关。方法:回顾性研究使用多中心RSNA腹部创伤CT数据集(2019–2023),涵盖23个中心的扫描。比较两种基础模型(MedCLIP,零样本;RadDINO,线性探测)与三种专用方法(CNN、Transformer、集成)。模型在3,147名患者(肠损伤患病率2.3%)上训练,于100例扩充测试集上评估。为分离阴性类影响,分别在无肠损伤但伴实质器官损伤(n=58)与无腹部病理者(n=50)中评估特异性。结果:基础模型判别能力与专用模型相当(AUC 0.64–0.68 vs 0.58–0.64),敏感性更高(79%–91% vs 41%–74%),但特异性更低(33%–50% vs 50%–88%)。所有模型在无腹部病理者中特异性均较高(84%–100%)。当存在实质器官损伤时,基础模型特异性显著下降(50–51百分点),而专用模型仅下降12–41百分点。结论:基础模型虽无需任务特定训练即可匹配判别性能,但其特异性缺陷主要由阴性类别的混杂异质性引起,而非单纯低患病率。对标注训练的依赖程度越高,对负样本异质性的抗性越强,提示临床应用前需进行适配。
原文摘要 · Abstract (English)
Purpose: Translating foundation models into clinical practice requires evaluating their performance under compound distribution shift, where severe class imbalance coexists with heterogeneous imaging appearances. This challenge is relevant for traumatic bowel injury, a rare but high-mortality diagnosis. We investigated whether specificity deficits in foundation models are associated with heterogeneity in the negative class. Methods: This retrospective study used the multi-institutional, RSNA Abdominal Traumatic Injury CT dataset (2019-2023), comprising scans from 23 centres. Two foundation models (MedCLIP, zero-shot; RadDINO, linear probe) were compared against three task-specific approaches (CNN, Transformer, Ensemble). Models were trained on 3,147 patients (2.3% bowel injury prevalence) and evaluated on an enriched 100-patient test set. To isolate negative-class effects, specificity was assessed in patients without bowel injury who had concurrent solid organ injury (n=58) versus no abdominal pathology (n=50). Results: Foundation models achieved equivalent discrimination to task-specific models (AUC, 0.64-0.68 versus 0.58-0.64) with higher sensitivity (79-91% vs 41-74%) but lower specificity (33-50% vs 50-88%). All models demonstrated high specificity in patients without abdominal pathology (84-100%). When solid organ injuries were present, specificity declined substantially for foundation models (50-51 percentage points) compared with smaller reductions of 12-41 percentage points for task-specific models. Conclusion: Foundation models matched task-specific discrimination without task-specific training, but their specificity deficits were driven primarily by confounding negative-class heterogeneity rather than prevalence alone. Susceptibility to negative-class heterogeneity decreased progressively with labelled training, suggesting adaptation is required before clinical implementation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。