用报告中的影像发现构建可审计的CT表型识别模型,发现模型常依赖无关特征。
Auditable CT Phenotyping Through Report-derived Radiological Observations

- 基于报告提取37万+影像观察,训练可追踪决策路径的模型ACT
- 零样本和线性探测下均优于现有方法,但仅97个观察主导221个表型预测
- 能识别并干预模型误用的非有效证据,适合临床验证与可信医疗AI研究
医学影像基础模型可从计算机断层扫描(CT)中预测临床表型,但高性能模型是否真正读取疾病特异性发现仍存疑。我们通过基于报告衍生影像观察构建的可审计CT表型识别(ACT)模型,在221个电子健康记录(EHR)表型上进行测试。模型在38,317名患者数据上训练,挖掘出376,194个影像观察,并在25,183名未见患者上评估。ACT在零样本标注下超越五种视觉-语言基线,且在未见过的胸部增强CT图像上,零样本评分(0.651对0.572)和线性探测评分(0.709对0.662)均更优。分析显示,仅有97个观察占据221个表型的首位,其中一句描述主动脉和冠状动脉钙化的短语在20个表型中排名第一,包括骨质疏松、尿路感染和重度抑郁障碍。将观察库限制为临床医生指定证据后,86个表型的预测探针转向相关观察,且准确率无损(0.751对0.741)。因此,当前准确的基于CT的EHR表型识别可能依赖于无效证据,而ACT可识别并干预此类偏差。
原文摘要 · Abstract (English)
Medical image foundation models can predict clinical phenotypes from computed tomography (CT), but strong performance leaves open whether they read disease-specific findings or shortcuts that correlate with the diagnosis. We tested this in 221 electronic-health-record (EHR) phenotypes using Auditable CT phenotyping (ACT), built on report-derived radiological observations. We trained ACT on 38,317 patients, mined 376,194 observations and evaluated it in 25,183 held-out patients. ACT exceeded five vision-language baselines on zero-shot annotation, and CT-CLIP across 221 phenotypes from unseen CT pulmonary angiography, both under zero-shot scoring (0.651 versus 0.572) and under linear probing (0.709 versus 0.662). Reading each probe exposes what accuracy conceals: only 97 observations occupy the 221 rank-1 positions, and one phrase describing aortic and coronary calcification ranks first for 20 phenotypes, including osteoporosis, urinary tract infection and major depressive disorder. Restricting the bank to clinician-specified evidence redirects those probes onto phenotype-related observations in 86 phenotypes at no accuracy cost (0.751 versus 0.741). Accurate CT-based EHR phenotyping can therefore rest on observations that are not valid evidence for the coded phenotype and that ACT can identify and intervene on.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。