arXiv:2507.03908cs.CV2025-07

用最优传输对齐影像与病灶标签,让报告更准确

Bridging Vision and Language: Optimal Transport-Driven Radiology Report Generation via LLMs

  • 用最优传输对齐图像特征与报告中的病灶标签
  • 在MIMIC-CXR和IU X-Ray上同时提升语言与临床表现
  • 适合需要高临床可信度的医学AI报告生成场景

放射科报告生成是医疗AI的重要应用,虽已取得显著进展,但通用大模型往往偏重语言流畅性而忽视临床有效性,难以捕捉影像与文本间的关联,导致实用性不足。为此,我们提出基于最优传输的放射科报告生成框架OTDRG,利用最优传输(OT)对齐图像视觉特征与报告中提取的病灶标签,有效弥合跨模态鸿沟。核心为对齐与微调模块,通过编码标签与图像特征,最小化跨模态距离,并融合双模态特征用于大模型微调。此外,设计新颖的疾病预测模块,在验证与测试阶段预测图像中的病灶标签。在MIMIC-CXR和IU X-Ray数据集上,OTDRG在自然语言生成(NLG)与临床效能(CE)指标上均达到当前最优水平,生成报告兼具语言连贯性与临床准确性。

原文摘要 · Abstract (English)

Radiology report generation represents a significant application within medical AI, and has achieved impressive results. Concurrently, large language models (LLMs) have demonstrated remarkable performance across various domains. However, empirical validation indicates that general LLMs tend to focus more on linguistic fluency rather than clinical effectiveness, and lack the ability to effectively capture the relationship between X-ray images and their corresponding texts, thus resulting in poor clinical practicability. To address these challenges, we propose Optimal Transport-Driven Radiology Report Generation (OTDRG), a novel framework that leverages Optimal Transport (OT) to align image features with disease labels extracted from reports, effectively bridging the cross-modal gap. The core component of OTDRG is Alignment \& Fine-Tuning, where OT utilizes results from the encoding of label features and image visual features to minimize cross-modal distances, then integrating image and text features for LLMs fine-tuning. Additionally, we design a novel disease prediction module to predict disease labels contained in X-ray images during validation and testing. Evaluated on the MIMIC-CXR and IU X-Ray datasets, OTDRG achieves state-of-the-art performance in both natural language generation (NLG) and clinical efficacy (CE) metrics, delivering reports that are not only linguistically coherent but also clinically accurate.

医学报告生成最优传输跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。