通过结构观察对比学习,提升CT报告生成的准确性与效率。
Structure Observation Driven Image-Text Contrastive Learning for Computed Tomography Report Generation
- 用可学习的结构查询捕捉CT图像中的解剖结构,实现图像与文本的结构级对齐。
- 在两个公开数据集上达到新最好效果,报告生成准确率显著提升。
- 适合医学AI研究者和临床辅助系统开发者参考使用。
计算机断层扫描报告生成(CTRG)旨在自动化放射科报告流程,减轻报告撰写负担并提升患者护理效率。尽管深度学习在X光报告生成中取得显著进展,但其在CTRG中的表现可能受限于CT图像数据量大、描述细节复杂等问题。本文提出一种两阶段(结构学习与报告生成)框架,采用有效的结构级图像-文本对比学习。第一阶段,一组可学习的结构特定视觉查询观察CT图像中的对应结构,生成的观察标记与伴随放射科报告中提取的结构特定文本特征进行结构级图像-文本对比损失优化。此外,提出基于文本-文本相似性的软伪标签,缓解非配对图像与报告间语义相同但未配对导致的假阴性问题。从而模型学会图像与报告间的结构级语义对应关系。进一步提出动态多样性增强负样本队列,引导网络区分各类异常。第二阶段,冻结视觉结构查询,用于选择关键图像块嵌入以表征每个解剖结构,减少无关区域干扰并降低内存消耗。同时加入文本解码器并训练生成报告。在两个公开数据集上的大量实验表明,该框架在临床效率方面建立了新的最先进性能,各组件均有效。
原文摘要 · Abstract (English)
Computed Tomography Report Generation (CTRG) aims to automate the clinical radiology reporting process, thereby reducing the workload of report writing and facilitating patient care. While deep learning approaches have achieved remarkable advances in X-ray report generation, their effectiveness may be limited in CTRG due to larger data volumes of CT images and more intricate details required to describe them. This work introduces a novel two-stage (structure- and report-learning) framework tailored for CTRG featuring effective structure-wise image-text contrasting. In the first stage, a set of learnable structure-specific visual queries observe corresponding structures in a CT image. The resulting observation tokens are contrasted with structure-specific textual features extracted from the accompanying radiology report with a structure-wise image-text contrastive loss. In addition, text-text similarity-based soft pseudo targets are proposed to mitigate the impact of false negatives, i.e., semantically identical image structures and texts from non-paired images and reports. Thus, the model learns structure-level semantic correspondences between CT images and reports. Further, a dynamic, diversity-enhanced negative queue is proposed to guide the network in learning to discriminate various abnormalities. In the second stage, the visual structure queries are frozen and used to select the critical image patch embeddings depicting each anatomical structure, minimizing distractions from irrelevant areas while reducing memory consumption. Also, a text decoder is added and trained for report generation.Our extensive experiments on two public datasets demonstrate that our framework establishes new state-of-the-art performance for CTRG in clinical efficiency, and its components are effective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。