arXiv:2608.00147cs.CVcs.LG2026-08

让医学影像理解按临床概念分层,提升可解释性与准确性。

RadPRISM: Schema-stratified radiology-report supervision for concept-disentangled image representations and visual grounding

论文配图:RadPRISM: Schema-stratified radiology-report supervision for concept-disentangled image representations and visual grounding
图 1 · 摘自论文原文
  • 按临床概念分空间对齐,实现概念解耦的视觉表示
  • 在胸部X光上零样本分类准确率从71.7%提升至86.8%
  • 医生评测显示概念检索正确率78%,适合临床可解释场景

视觉-语言预训练从放射科报告中学习丰富的医学图像表征,但以往模型多在单一共享嵌入空间中运行,导致概念级结构和可解释性需事后恢复,限制了模型透明度与临床应用。我们提出RadPRISM,将临床定义的放射科模板作为分层轴:本地大语言模型从自由文本报告中提取各概念的文本片段,并将每个临床概念映射到专属的视觉子空间,使概念分层成为直接、顶层的对齐监督。在包含203,602例检查的内部多年度数据库上,基于19个概念的胸片任务中,RadPRISM将内部数据集零样本分类的宏平均AUC从0.717(95%置信区间0.710–0.723)提升至0.868(95%置信区间0.863–0.872),优于匹配的全局对齐基线;在外部零样本分类表现与专用模型CARZero相当,但在点位游戏视觉定位任务中最高超出其4.3倍。此外,放射科医生评估显示,概念分层检索在前3名内的宏平均正确率达0.78,展现出报告级检索与固定标签词汇无法表达的解耦描述发现。RadPRISM生成具有区分性、空间忠实性且天然概念分层的表征,由临床专家定义并透明可查。

原文摘要 · Abstract (English)

Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a single shared embedding space, so concept-level structure and interpretability must be recovered post hoc, limiting model transparency and, hence, clinical utility. We introduce RadPRISM, which makes a clinician-defined radiology schema a designated stratification axis: an on-premise large language model extracts per-concept text spans from free-text reports, and each clinical concept is aligned in its own dedicated visual subspace, turning concept stratification into direct, top-level alignment supervision. Instantiated on chest radiographs with a 19-concept schema over $203{,}602$ examinations from an internal multi-year archive, RadPRISM improved internal dataset zero-shot classification from $0.717$ (95% CI, $0.710-0.723$) to $0.868$ (95% CI, $0.863-0.872$) macro AUROC over a matched global-alignment baseline, performed on par with the purpose-built CARZero reference in external zero-shot classification while substantially outperforming it (up to 4.3-fold) in pointing-game visual grounding. In addition, a radiologist reader study demonstrated concept-stratified retrieval ability ($0.78$ macro retrieval correctness rate within rank 3), surfacing disentangled descriptive findings that report-level retrieval and fixed-label vocabularies cannot express. RadPRISM yields discriminative, spatially faithful, natively concept-stratified representations shaped by and transparently inspectable by clinicians.

医学影像概念解耦视觉定位可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。