arXiv:2501.12356cs.CV2025-01被引 7

用视觉语言模型自动生成胸部X光报告,效果优于传统方法

Vision-Language Models for Automated Chest X-ray Interpretation: Leveraging ViT and GPT-2

  • 融合ViT与GPT-2的多模态架构生成报告
  • SWIN-BART组合在ROUGE、BLEU、BERTScore上表现最优
  • 适合医学AI研究者和临床辅助系统开发者

放射学在现代医学中至关重要,因其非侵入性诊断能力。然而,手动生成非结构化医学报告耗时且易出错,成为临床流程中的瓶颈。尽管人工智能生成放射科报告已有进展,但实现详细准确的报告仍具挑战。本研究评估了多种融合计算机视觉与自然语言处理的多模态模型组合,用于生成全面的放射科报告。采用预训练的Vision Transformer (ViT-B16) 和 SWIN Transformer 作为图像编码器,BART 和 GPT-2 作为文本解码器。基于 IU-Xray 数据集的胸部X光图像与报告,评估了 SWIN-BART、SWIN-GPT-2、ViT-B16-BART 和 ViT-B16-GPT-2 四种模型的表现。结果表明,SWIN-BART 模型在所有评估指标(如 ROUGE、BLEU、BERTScore)上均表现最佳。

原文摘要 · Abstract (English)

Radiology plays a pivotal role in modern medicine due to its non-invasive diagnostic capabilities. However, the manual generation of unstructured medical reports is time consuming and prone to errors. It creates a significant bottleneck in clinical workflows. Despite advancements in AI-generated radiology reports, challenges remain in achieving detailed and accurate report generation. In this study we have evaluated different combinations of multimodal models that integrate Computer Vision and Natural Language Processing to generate comprehensive radiology reports. We employed a pretrained Vision Transformer (ViT-B16) and a SWIN Transformer as the image encoders. The BART and GPT-2 models serve as the textual decoders. We used Chest X-ray images and reports from the IU-Xray dataset to evaluate the usability of the SWIN Transformer-BART, SWIN Transformer-GPT-2, ViT-B16-BART and ViT-B16-GPT-2 models for report generation. We aimed at finding the best combination among the models. The SWIN-BART model performs as the best-performing model among the four models achieving remarkable results in almost all the evaluation metrics like ROUGE, BLEU and BERTScore.

医学影像多模态报告生成ViT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。