通过病灶原型对齐影像与报告,提升放射科报告生成的临床一致性。
Learning to See Locally and Align Clinically with Pathology Semantics for Radiology Report Generation

- 用共享病灶原型替代直接图像-报告配对,实现更自然的语义对齐。
- 在MIMIC-CXR等数据集上,报告生成与异常检测均显著提升。
- 适合需要高临床可信度的医学影像报告生成场景。
近期针对放射科的视觉-语言模型在标准报告生成基准上表现良好,但其鲁棒性和泛化能力受限于视觉与文本特征间不完善的对齐。现有方法或通过自回归报告监督隐式连接,或通过对比学习显式对齐。然而,仅靠自回归监督难以建立可靠的图文对齐,而对比学习可能将描述相似病灶但未配对的报告错误分离。这在放射科中尤为严重,因不同报告可能共享兼容的病灶语义而非真实负例。导致学习的表征无法围绕共同病灶概念组织,使解码器依赖预训练语言先验,生成看似合理却缺乏影像支持的报告。为此,我们提出PALM,一种病灶感知对齐框架。不同于直接匹配每对图像-报告并分离其余,PALM通过共享病灶原型对齐视觉与文本特征,为影像证据与文本发现提供临床有意义的桥梁,使具有相似病灶语义的病例趋向共同概念,而不分离兼容案例。此外,引入掩码证据建模,通过学习被遮蔽区域引起的语义变化,增强图像编码器对局部影像证据的敏感性。在MIMIC-CXR、IU X-Ray和MIMIC-ABN上的实验表明,PALM在报告生成和异常聚焦鲁棒性方面均有持续提升。
原文摘要 · Abstract (English)
Recent radiology-adapted vision-language models have achieved strong performance on standard report generation benchmarks, yet their robustness and generalization remain constrained by imperfect alignment and correlation between visual and textual features. Existing methods connect image and text either implicitly through autoregressive report supervision or explicitly through contrastive learning. However, autoregressive supervision alone is insufficient to establish reliable image-text alignment, while contrastive learning can push apart unpaired reports that describe related pathologies simply because they are not paired with the same image. This is problematic in radiology, where different reports may share compatible pathology semantics rather than being true negatives. As a result, the learned representation may fail to organize images and reports around shared pathology concepts, causing the decoder to rely on pretrained language priors and generate clinically plausible reports that are not fully supported by radiographic evidence. To address this issue, we propose PALM, a pathology-aware alignment framework for radiology report generation. Instead of directly matching each image-report pair while separating all others, PALM aligns visual and textual features through shared pathology prototypes. These prototypes provide a clinically meaningful bridge between radiographic evidence and textual findings, allowing cases with similar pathology semantics to move toward common concepts without separating compatible cases. In addition, we introduce Masked Evidence Modeling to strengthen the image encoder sensitivity to local radiographic evidence by learning semantic changes caused by masked image regions. Experiments on MIMIC-CXR, IU X-Ray, and MIMIC-ABN show that PALM consistently improves both report generation and abnormality-focused robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。