arXiv:2604.27559cs.CVcs.AI2026-04中稿 · Journal of Biomedi…被引 1

通过分层对齐图像与报告结构,提升放射科报告生成精度

RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation

论文配图:RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation
图 1 · 摘自论文原文
  • 构建图像与报告在段、句、词三级的跨模态对齐机制
  • 在IU-Xray和MIMIC-CXR上显著优于现有模型,报告生成更准确
  • 适合医学影像与自然语言生成交叉研究者参考

放射科报告生成(RRG)旨在通过自动从医学影像生成诊断报告,减轻放射科医生的工作负担并降低人为错误。其核心挑战在于实现复杂视觉特征与长篇报告层级结构之间的细粒度对齐。尽管近期方法提升了图像-文本表征学习能力,但大多将报告视为扁平序列,忽略了其分段与语义层次结构,导致跨模态对齐不精准,影响生成效果。为此,本文提出RIHA(Report-Image Hierarchical Alignment Transformer),一种端到端框架,在段落、句子、词汇三个层级实现影像与报告的多级对齐。该方法引入视觉特征金字塔(VFP)提取多尺度视觉特征,文本特征金字塔(TFP)表示多粒度文本结构,并通过跨模态层级对齐(CHA)模块结合最优传输算法,实现各层级的高效对齐。此外,解码器中引入相对位置编码(RPE),建模标记间的空间与语义关系,强化视觉特征与生成文本的细粒度匹配。在两个基准胸部X光数据集IU-Xray和MIMIC-CXR上的大量实验表明,RIHA在自然语言生成与临床有效性指标上均超越现有最先进模型。

原文摘要 · Abstract (English)

Radiology report generation (RRG) has emerged as a promising approach to alleviate radiologists' workload and reduce human errors by automatically generating diagnostic reports from medical images. A key challenge in RRG is achieving fine-grained alignment between complex visual features and the hierarchical structure of long-form radiology reports. Although recent methods have improved image-text representation learning, they often treat reports as flat sequences, overlooking their structured sections and semantic hierarchies. This simplification hinders precise cross-modal alignment and weakens RRG accuracy. To address this challenge, we propose RIHA (Report-Image Hierarchical Alignment Transformer), a novel end-to-end framework that performs multi-level alignment between radiological images and their corresponding reports across paragraph, sentence, and word levels. This hierarchical alignment enables more precise cross-modal mapping, essential for capturing the nuanced semantics embedded in clinical narratives. Specifically, RIHA introduces a Visual Feature Pyramid (VFP) to extract multi-scale visual features and a Text Feature Pyramid (TFP) to represent multi-granularity textual structures. These components are integrated through a Cross-modal Hierarchical Alignment (CHA) module, leveraging optimal transport to effectively align visual and textual features across various levels. Furthermore, we incorporate Relative Positional Encoding (RPE) into the decoder to model spatial and semantic relationships among tokens, enhancing the token-level alignment between visual features and generated text. Extensive experiments on two benchmark chest X-ray datasets, IU-Xray and MIMIC-CXR, demonstrate that RIHA outperforms existing state-of-the-art models in both natural language generation and clinical efficacy metrics.

报告生成医学影像跨模态对齐Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。