将CT基础模型蒸馏为可编辑的概念瓶颈,提升肺结节恶性预测的可解释性。
Distilling CT Foundation Models into Editable Concept Bottlenecks for Lung Nodule Malignancy Prediction

- 用放射科医生定义的8个特征构建概念瓶颈,从CT基础模型中提取可解释特征。
- 模型在内部测试集上达到0.86的AUROC,外部验证集达0.73,接近仅用结节大小的效果。
- 支持对概念进行可控干预和贡献分解,实现透明且可调节的恶性预测。
基础模型提供可迁移的CT表征,但直接基于这些嵌入的预测难以解释。本文构建了概念瓶颈模型,将两个冻结的CT基础模型表征映射到8个放射科医生定义的肺结节属性,并结合结节大小预测恶性程度。模型包括使用96³体素结节中心补丁的全CT自监督编码器CT-FM,以及使用50毫米裁剪区域的结节聚焦对比编码器FMCIB。在2,610个LIDC-IDRI结节上训练了8个岭回归概念头。恶性预测模型在LUNA25上训练,于保留的内部测试集及外部DLCS队列上评估。概念保真度通过五折交叉验证R²评估,恶性判别通过带95%置信区间的AUROC评估(采用患者分组自助抽样)。结果显示,FMCIB在模糊性(R² 0.24 vs. 0.11)、毛刺(0.17 vs. 0.08)、纹理(0.17 vs. 0.07)和分叶(0.15 vs. 0.05)上的概念保真度高于CT-FM。内部测试中,CT-FM与FMCIB的组合模型分别取得0.86(95% CI: 0.80–0.92)和0.86(0.79–0.92)的AUROC;外部验证中分别为0.72(0.68–0.75)和0.73(0.70–0.76),优于仅用结节大小(0.73)或对应嵌入探针(0.60 和 0.67)。预测结果可分解为特征级贡献并经概念干预修改。概念瓶颈实现了与结节大小相当的恶性判别能力,同时提供透明解释;不同基础模型表征下概念恢复效果差异表明,概念提取依赖于底层表征质量。
原文摘要 · Abstract (English)
Foundation models provide transferable CT representations, but predictions based directly on these embeddings are difficult to interpret. We developed concept bottleneck models that map two frozen CT foundation-model representations to eight radiologist-defined pulmonary-nodule attributes and predict malignancy from the estimated concepts and nodule size. The models included CT-FM, a whole-CT self-supervised encoder using a 96^3-voxel nodule-centered patch, and FMCIB, a nodule-focused contrastive encoder using a 50-mm crop. Eight ridge-regression concept heads were trained on 2,610 LIDC-IDRI nodules. Malignancy models were trained on LUNA25 and evaluated on a held-out internal test set and the external DLCS cohort. Concept fidelity was assessed using five-fold cross-validated R^2, and malignancy discrimination was assessed using AUROC with 95% confidence intervals estimated by patient-grouped bootstrap resampling. Concept fidelity was modest but higher for FMCIB than CT-FM for subtlety (R2, 0.24 vs. 0.11), spiculation (0.17 vs. 0.08), texture (0.17 vs. 0.07), and lobulation (0.15 vs. 0.05). Internally, the CT-FM and FMCIB concept+size models achieved AUROCs of 0.86 (95% CI, 0.80-0.92) and 0.86 (0.79-0.92), respectively. Externally, AUROCs were 0.72 (0.68-0.75) and 0.73 (0.70-0.76), compared with 0.73 for nodule size alone and 0.60 and 0.67 for the corresponding embedding only probes. Additive predictions could be decomposed into feature-level contributions and modified through controlled concept interventions. Concept bottlenecks provided transparent malignancy predictions with discrimination similar to nodule size alone, while differences in concept fidelity suggest that concept recovery depends on the underlying foundation-model representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。