用解剖结构层次知识增强CT影像与报告的对齐,提升病灶识别能力。
Learning Anatomy-Grounded CT Vision-Language Representations with Organ-Hierarchical Report Knowledge

- 基于器官层级知识构建报告结构,指导细粒度图像-文本对齐
- 零样本诊断AUC达84.9(CT-RATE)和72.2(RAD-ChestCT),优于现有方法
- 适合医学视觉语言预训练、肺部疾病检测等临床场景研究者
医学视觉语言预训练(VLP)通过配对的CT影像与放射科报告实现可扩展的表征学习,但现有方法多将整幅扫描与全文报告对齐,或局部图像区域与文本片段对齐。这类方法未充分利用放射科报告的关键特性:病变按解剖结构组织,异常描述包含器官、疾病概念、位置及严重程度属性。本文提出OKA-CT,一种基于器官层级知识增强的CT-报告VLP框架。OKA-CT首先利用放射科报告解析与大模型辅助语义结构化,将自由文本报告转化为器官条件化的知识。该提取的层级结构贯穿两个学习阶段:第一阶段通过细粒度器官条件监督,将解剖定位证据注入CT视觉表征;第二阶段使用器官特异性报告证据,引导结构化报告-CT对比学习,其中由层级导出的语义软目标将具有共享器官级发现的非配对病例视为弱语义正例而非统一负例。此外,轻量级查询式全局分支进一步聚合与疾病相关的体积分量证据,生成全扫描表征。在CT-RATE与RAD-ChestCT数据集上,OKA-CT实现零样本异常诊断的AUROC分别为84.9与72.2,超越先前基线。检索与补丁遮蔽分析进一步表明其报告-图像对齐更优,且对疾病相关解剖区域更敏感。
原文摘要 · Abstract (English)
Medical vision-language pretraining (VLP) from paired CT images and radiology reports enables scalable representation learning, but most existing methods align either whole scans with entire reports or local image regions with text fragments. These formulations underuse a key property of radiology reports: findings are organized around anatomical structures, with abnormalities described by organs, disease concepts, locations, and severity-related attributes. We propose OKA-CT, an organ-hierarchical knowledge-augmented framework for CT-report VLP. OKA-CT first converts free-text reports into organ-conditioned knowledge using radiology report parsing and LLM-assisted semantic structuring. The extracted hierarchy is used across two learning stages. Stage~1 injects anatomy-grounded evidence into the CT visual representation through fine-grained organ-conditioned supervision, while Stage~2 uses organ-specific report evidence to guide structured report-CT contrastive learning, where hierarchy-derived semantic soft targets treat non-paired cases with shared organ-level findings as weak semantic positives rather than uniform negatives. A lightweight query-based global branch further aggregates disease-relevant volumetric evidence for whole-scan representation. On CT-RATE and RAD-ChestCT datasets, OKA-CT achieves zero-shot abnormality diagnosis AUROCs of 84.9 and 72.2, outperforming prior CT VLP baselines. Retrieval and patch-occlusion analyses further show improved report-image alignment and stronger sensitivity to disease-associated anatomical regions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。