用视觉变压器+跨模态注意力生成更准的胸部X光描述。
Advanced Chest X-Ray Analysis via Transformer-Based Image Descriptors and Cross-Model Attention Mechanism
- 用ViT提取图像特征,通过跨模态注意力融合文本信息。
- 在NIH数据集上多项指标领先,最高CIDEr达0.857。
- 适合医疗辅助诊断,提升放射科医生效率。
胸部X光检查是检测多种胸腔疾病的关键环节。本文提出一种新图像描述生成模型,结合视觉变换器(ViT)编码器、跨模态注意力机制与基于GPT-4的解码器。ViT从胸部X光中捕捉高质量视觉特征,通过跨模态注意力与文本数据融合,提升描述的准确性、上下文连贯性与丰富度;GPT-4解码器将融合特征转化为精准且相关的图像标题。模型在国家卫生研究院(NIH)和印第安纳大学(IU)胸部X光数据集上测试。在IU数据集上,得分分别为:B-1 0.854,CIDEr 0.883,METEOR 0.759,ROUGE-L 0.712。在NIH数据集上所有指标均最优:BLEU 1–4(0.825, 0.788, 0.765, 0.752),CIDEr 0.857,METEOR 0.726,ROUGE-L 0.705。该框架有潜力提升胸部X光评估效果,辅助放射科医生实现更精确高效的诊断。
原文摘要 · Abstract (English)
The examination of chest X-ray images is a crucial component in detecting various thoracic illnesses. This study introduces a new image description generation model that integrates a Vision Transformer (ViT) encoder with cross-modal attention and a GPT-4-based transformer decoder. The ViT captures high-quality visual features from chest X-rays, which are fused with text data through cross-modal attention to improve the accuracy, context, and richness of image descriptions. The GPT-4 decoder transforms these fused features into accurate and relevant captions. The model was tested on the National Institutes of Health (NIH) and Indiana University (IU) Chest X-ray datasets. On the IU dataset, it achieved scores of 0.854 (B-1), 0.883 (CIDEr), 0.759 (METEOR), and 0.712 (ROUGE-L). On the NIH dataset, it achieved the best performance on all metrics: BLEU 1--4 (0.825, 0.788, 0.765, 0.752), CIDEr (0.857), METEOR (0.726), and ROUGE-L (0.705). This framework has the potential to enhance chest X-ray evaluation, assisting radiologists in more precise and efficient diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。