提出双阶段框架,让肺部X光报告生成更关注病灶特征。
A Disease-Aware Dual-Stage Framework for Chest X-ray Report Generation
- 第一阶段用交叉注意力学疾病语义令牌,第二阶段融合疾病与视觉特征。
- 在三个数据集上报告临床准确率和语言质量均达领先水平。
- 适合医疗AI研究者,尤其关注医学影像生成的可解释性。
从胸部X光片生成放射科报告是人工智能的重要任务,有望显著减轻放射科医生的工作负担并缩短患者等待时间。尽管近期取得进展,现有方法在视觉表征中缺乏足够的疾病感知能力,且视觉-语言对齐不足,难以满足医学图像分析的专业需求。因此,这些模型常忽略胸部X光中的关键病理特征,生成的报告临床准确性不高。为解决上述问题,本文提出一种新型双阶段疾病感知框架。第一阶段通过交叉注意力机制和多标签分类学习对应特定病理类别的疾病感知语义令牌(DASTs),同时利用对比学习对齐视觉与语言表示。第二阶段引入疾病-视觉注意力融合模块(DVAF)整合疾病感知表示与视觉特征,并设计双模态相似性检索机制(DMSR),结合视觉与疾病特异性相似性检索相关范例,为报告生成提供上下文指导。在基准数据集(CheXpert Plus、IU X-ray、MIMIC-CXR)上的大量实验表明,该疾病感知框架在胸部X光报告生成任务中达到当前最优性能,临床准确性和语言质量均有显著提升。
原文摘要 · Abstract (English)
Radiology report generation from chest X-rays is an important task in artificial intelligence with the potential to greatly reduce radiologists' workload and shorten patient wait times. Despite recent advances, existing approaches often lack sufficient disease-awareness in visual representations and adequate vision-language alignment to meet the specialized requirements of medical image analysis. As a result, these models usually overlook critical pathological features on chest X-rays and struggle to generate clinically accurate reports. To address these limitations, we propose a novel dual-stage disease-aware framework for chest X-ray report generation. In Stage~1, our model learns Disease-Aware Semantic Tokens (DASTs) corresponding to specific pathology categories through cross-attention mechanisms and multi-label classification, while simultaneously aligning vision and language representations via contrastive learning. In Stage~2, we introduce a Disease-Visual Attention Fusion (DVAF) module to integrate disease-aware representations with visual features, along with a Dual-Modal Similarity Retrieval (DMSR) mechanism that combines visual and disease-specific similarities to retrieve relevant exemplars, providing contextual guidance during report generation. Extensive experiments on benchmark datasets (i.e., CheXpert Plus, IU X-ray, and MIMIC-CXR) demonstrate that our disease-aware framework achieves state-of-the-art performance in chest X-ray report generation, with significant improvements in clinical accuracy and linguistic quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。