让视觉语言模型学会识别胸片变化,提升医学报告理解能力
CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding
- 用大模型解析报告,提取时间演变结构和细粒度标注
- 提出对比性描述模型CoCa-CXR,准确识别胸片前后差异
- 在肺部病变进展分类上比现有最优模型高4.8%,适合医疗影像分析研究者
视觉语言模型在医学图像分析中表现优异,因其能从图像和报告中学习丰富语义。以往工作主要关注图像与文本表示的对齐以增强图像理解,但胸片报告中常涉及前后图像的对比,如何将进展描述与图像间的语义差异对齐仍研究不足。本文提出两个组件:(1) 基于大语言模型的胸片报告处理流程,可分离描述与比较上下文,并提取细粒度标注;(2) 针对胸片的对比性描述模型CoCa-CXR,能够同时描述图像及其时间进展。CoCa-CXR引入新型区域交叉注意力模块,精准定位成对胸片的局部差异。大量实验表明,该模型在进展分析与报告生成任务上均优于先前方法。在MS-CXR-T进展分类任务中,五类肺部疾病平均测试准确率达65.0%,较之前最优模型BioViL-T提升4.8%。在MIMIC-CXR上的RadGraph F1达24.2%,接近Med-Gemini基础模型水平。
原文摘要 · Abstract (English)
Vision-language models have proven to be of great benefit for medical image analysis since they learn rich semantics from both images and reports. Prior efforts have focused on better alignment of image and text representations to enhance image understanding. However, though explicit reference to a prior image is common in Chest X-Ray (CXR) reports, aligning progression descriptions with the semantics differences in image pairs remains under-explored. In this work, we propose two components to address this issue. (1) A CXR report processing pipeline to extract temporal structure. It processes reports with a large language model (LLM) to separate the description and comparison contexts, and extracts fine-grained annotations from reports. (2) A contrastive captioner model for CXR, namely CoCa-CXR, to learn how to both describe images and their temporal progressions. CoCa-CXR incorporates a novel regional cross-attention module to identify local differences between paired CXR images. Extensive experiments show the superiority of CoCa-CXR on both progression analysis and report generation compared to previous methods. Notably, on MS-CXR-T progression classification, CoCa-CXR obtains 65.0% average testing accuracy on five pulmonary conditions, outperforming the previous state-of-the-art (SOTA) model BioViL-T by 4.8%. It also achieves a RadGraph F1 of 24.2% on MIMIC-CXR, which is comparable to the Med-Gemini foundation model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。