arXiv:2505.07689cs.CV2025-05

通过解剖结构对齐提升医学影像报告生成的准确性

Anatomical Attention Alignment representation for Radiology Report Generation

  • 将解剖结构知识库与图像块特征融合,构建结构化视觉表示
  • 在IU X-Ray和MIMIC-CXR数据集上显著提升报告生成质量
  • 适合医疗AI研究者及临床辅助系统开发者使用

自动化放射科报告生成(RRG)旨在生成医学图像的详细描述,减轻放射科医生的工作负担并提高高质量诊断服务的可及性。现有编码器-解码器模型仅依赖原始输入图像提取的视觉特征,难以理解空间结构与语义关系,常导致文本生成效果不佳。为此,我们提出解剖注意力对齐网络(A3Net),通过构建超视觉表示来增强视觉-文本理解。该方法将解剖结构知识字典与图像块级视觉特征结合,使模型能有效关联图像区域与其对应的解剖实体。这种结构化表示提升了语义推理能力、可解释性及跨模态对齐效果,最终增强了生成报告的准确性和临床相关性。在IU X-Ray和MIMIC-CXR数据集上的实验结果表明,A3Net显著改善了视觉感知与文本生成质量。代码已开源。

原文摘要 · Abstract (English)

Automated Radiology report generation (RRG) aims at producing detailed descriptions of medical images, reducing radiologists' workload and improving access to high-quality diagnostic services. Existing encoder-decoder models only rely on visual features extracted from raw input images, which can limit the understanding of spatial structures and semantic relationships, often resulting in suboptimal text generation. To address this, we propose Anatomical Attention Alignment Network (A3Net), a framework that enhance visual-textual understanding by constructing hyper-visual representations. Our approach integrates a knowledge dictionary of anatomical structures with patch-level visual features, enabling the model to effectively associate image regions with their corresponding anatomical entities. This structured representation improves semantic reasoning, interpretability, and cross-modal alignment, ultimately enhancing the accuracy and clinical relevance of generated reports. Experimental results on IU X-Ray and MIMIC-CXR datasets demonstrate that A3Net significantly improves both visual perception and text generation quality. Our code is available at \href{https://github.com/Vinh-AI/A3Net}{GitHub}.

医学报告生成视觉-语言对齐解剖结构AI辅助诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。