通过可学习检索增强图像与报告对齐,提升放射科报告生成质量。
Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report Generation
- 引入可学习检索模块,基于语义层次动态匹配相关报告,增强罕见病图像表征。
- 在MIMIC-CXR上提升7.4%,在IU X-Ray上提升2.9%,显著改善跨模态融合效果。
- 适合医学影像生成、低频疾病识别等需要精准对齐的临床场景。
自动化放射科报告生成对提升诊断效率、减轻医疗人员负担至关重要。现有方法面临疾病类别不平衡和跨模态融合不足的挑战。为此,本文提出可学习检索增强的视觉-文本对齐与融合框架(REVTAF),有效解决类别不平衡与视觉-文本融合问题。该框架包含两个核心组件:(1) 可学习检索增强器(LRE),利用双曲空间中的语义层次和批内上下文信息,通过基于排名的度量实现自适应检索,增强输入图像表示,尤其改善尾部类别的表现;(2) 细粒度视觉-文本对齐与融合策略,确保多源交叉注意力图的一致性,并采用基于最优传输的交叉注意力机制,动态整合任务相关的文本知识以提升生成质量。结合自适应检索与多源对齐融合,REVTAF在弱图像-报告级监督下实现精细的跨模态集成,并有效缓解数据不平衡问题。实验表明,该方法在MIMIC-CXR数据集上平均提升7.4%,在IU X-Ray数据集上提升2.9%,优于当前主流多模态大模型(如GPT系列)。
原文摘要 · Abstract (English)
Automated radiology report generation is essential for improving diagnostic efficiency and reducing the workload of medical professionals. However, existing methods face significant challenges, such as disease class imbalance and insufficient cross-modal fusion. To address these issues, we propose the learnable Retrieval Enhanced Visual-Text Alignment and Fusion (REVTAF) framework, which effectively tackles both class imbalance and visual-text fusion in report generation. REVTAF incorporates two core components: (1) a Learnable Retrieval Enhancer (LRE) that utilizes semantic hierarchies from hyperbolic space and intra-batch context through a ranking-based metric. LRE adaptively retrieves the most relevant reference reports, enhancing image representations, particularly for underrepresented (tail) class inputs; and (2) a fine-grained visual-text alignment and fusion strategy that ensures consistency across multi-source cross-attention maps for precise alignment. This component further employs an optimal transport-based cross-attention mechanism to dynamically integrate task-relevant textual knowledge for improved report generation. By combining adaptive retrieval with multi-source alignment and fusion, REVTAF achieves fine-grained visual-text integration under weak image-report level supervision while effectively mitigating data imbalance issues. The experiments demonstrate that REVTAF outperforms state-of-the-art methods, achieving an average improvement of 7.4% on the MIMIC-CXR dataset and 2.9% on the IU X-Ray dataset. Comparisons with mainstream multimodal LLMs (e.g., GPT-series models), further highlight its superiority in radiology report generation https://github.com/banbooliang/REVTAF-RRG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。