用异常对齐的自举法提升肺栓塞影像诊断与报告生成
Abn-BLIP: Abnormality-aligned Bootstrapping Language-Image Pre-training for Pulmonary Embolism Diagnosis and Report Generation from CTPA
- 通过可学习查询和跨模态注意力对齐异常区域
- 减少漏诊,报告准确率与临床相关性显著提升
- 适合医学影像智能辅助诊断研究者参考
医学影像在现代医疗中至关重要,其中计算机断层扫描肺动脉造影(CTPA)是诊断肺栓塞及其他胸腔疾病的关键工具。然而,解读CTPA图像并生成准确放射科报告仍面临巨大挑战。本文提出Abn-BLIP(异常对齐自举语言-图像预训练),一种先进的诊断模型,旨在通过将异常发现精准对齐,提升报告的准确性和完整性。该模型利用可学习查询和跨模态注意力机制,在异常检测、减少漏诊以及结构化报告生成方面均优于现有方法。实验表明,Abn-BLIP在准确性和临床相关性上超越当前最先进的医疗视觉-语言模型及3D报告生成方法。结果表明,融合多模态学习策略在改善放射科报告方面具有巨大潜力。源代码见:https://github.com/zzs95/abn-blip。
原文摘要 · Abstract (English)
Medical imaging plays a pivotal role in modern healthcare, with computed tomography pulmonary angiography (CTPA) being a critical tool for diagnosing pulmonary embolism and other thoracic conditions. However, the complexity of interpreting CTPA scans and generating accurate radiology reports remains a significant challenge. This paper introduces Abn-BLIP (Abnormality-aligned Bootstrapping Language-Image Pretraining), an advanced diagnosis model designed to align abnormal findings to generate the accuracy and comprehensiveness of radiology reports. By leveraging learnable queries and cross-modal attention mechanisms, our model demonstrates superior performance in detecting abnormalities, reducing missed findings, and generating structured reports compared to existing methods. Our experiments show that Abn-BLIP outperforms state-of-the-art medical vision-language models and 3D report generation methods in both accuracy and clinical relevance. These results highlight the potential of integrating multimodal learning strategies for improving radiology reporting. The source code is available at https://github.com/zzs95/abn-blip.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。