arXiv:2502.20056cs.CVcs.AI2025-02CVPR被引 67

用多视角纵向X光片提升报告生成准确率

Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report Generation

  • 融合当前多视角图像与历史检查数据的对比学习
  • 在三个数据集上报告生成指标均优于现有方法
  • 适合需要追踪病情发展的医学AI研究者

自动化放射科报告生成可有效减轻放射科医生的工作负担。然而,现有方法大多仅依赖单视图或固定视角图像来建模当前病情,限制了诊断准确率并忽略了疾病进展。尽管部分方法使用纵向数据追踪病情变化,仍依赖单张图像分析当前就诊情况。为此,我们提出增强型多视角纵向对比学习用于胸部X光报告生成(MLRG)。具体地,引入一种融合当前多视角图像空间信息与纵向数据时间信息的对比学习方法,并利用放射科报告中固有的时空信息监督视觉与文本表征的预训练。此外,提出分词缺失编码技术,灵活处理患者特定先验知识缺失问题,使模型能基于可用信息生成更精准报告。在MIMIC-CXR、MIMIC-ABN和Two-view CXR数据集上的大量实验表明,本方法优于近期先进方法:在MIMIC-CXR上BLEU-4提升2.3%,在MIMIC-ABN上F1分数提升5.5%,在Two-view CXR上F1 RadGraph提升2.7%。

原文摘要 · Abstract (English)

Automated radiology report generation offers an effective solution to alleviate radiologists' workload. However, most existing methods focus primarily on single or fixed-view images to model current disease conditions, which limits diagnostic accuracy and overlooks disease progression. Although some approaches utilize longitudinal data to track disease progression, they still rely on single images to analyze current visits. To address these issues, we propose enhanced contrastive learning with Multi-view Longitudinal data to facilitate chest X-ray Report Generation, named MLRG. Specifically, we introduce a multi-view longitudinal contrastive learning method that integrates spatial information from current multi-view images and temporal information from longitudinal data. This method also utilizes the inherent spatiotemporal information of radiology reports to supervise the pre-training of visual and textual representations. Subsequently, we present a tokenized absence encoding technique to flexibly handle missing patient-specific prior knowledge, allowing the model to produce more accurate radiology reports based on available prior knowledge. Extensive experiments on MIMIC-CXR, MIMIC-ABN, and Two-view CXR datasets demonstrate that our MLRG outperforms recent state-of-the-art methods, achieving a 2.3% BLEU-4 improvement on MIMIC-CXR, a 5.5% F1 score improvement on MIMIC-ABN, and a 2.7% F1 RadGraph improvement on Two-view CXR.

报告生成纵向数据对比学习X光影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。