arXiv:2410.00379cs.CVcs.AI2024-10CVPR被引 45

针对胸部X光报告生成,构建新基准并提出多阶段预训练模型。

CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset

  • 采用自回归与图文对比学习的多阶段预训练策略。
  • 在CheXpert Plus数据集上显著优于现有模型,提升报告生成质量。
  • 适合医学AI研究者快速掌握该领域最新技术进展。

基于X光图像的医疗报告生成(MRG)是人工智能在医疗领域的重要方向,能有效减轻诊断负担和缩短患者等待时间。尽管取得显著进展,但受限于基准数据集不足及现有大模型在该专业领域的性能提升有限,任务发展面临瓶颈。特别是新发布的CheXpert Plus数据集缺乏对比评估算法及其结果,仅提供数据本身,导致后续算法的训练、评估与比较困难。为此,我们在CheXpert Plus数据集上全面评测了主流X光报告生成模型与大语言模型(LLMs)。我们提出的基准可为后续算法提供可靠的对比基础,并引导研究人员快速掌握该领域的最先进技术。更重要的是,我们提出一种面向X光图像报告生成的大模型,采用多阶段预训练策略,包括自监督自回归生成、图像-报告对比学习以及监督微调。大量实验表明,基于Mamba的自回归预训练能有效编码X光图像,图像-文本对比预训练进一步对齐特征空间,获得更优性能。源代码可在 <https://github.com/Event-AHU/Medical_Image_Analysis> 获取。

原文摘要 · Abstract (English)

X-ray image-based medical report generation (MRG) is a pivotal area in artificial intelligence which can significantly reduce diagnostic burdens and patient wait times. Despite significant progress, we believe that the task has reached a bottleneck due to the limited benchmark datasets and the existing large models' insufficient capability enhancements in this specialized domain. Specifically, the recently released CheXpert Plus dataset lacks comparative evaluation algorithms and their results, providing only the dataset itself. This situation makes the training, evaluation, and comparison of subsequent algorithms challenging. Thus, we conduct a comprehensive benchmarking of existing mainstream X-ray report generation models and large language models (LLMs), on the CheXpert Plus dataset. We believe that the proposed benchmark can provide a solid comparative basis for subsequent algorithms and serve as a guide for researchers to quickly grasp the state-of-the-art models in this field. More importantly, we propose a large model for the X-ray image report generation using a multi-stage pre-training strategy, including self-supervised autoregressive generation and Xray-report contrastive learning, and supervised fine-tuning. Extensive experimental results indicate that the autoregressive pre-training based on Mamba effectively encodes X-ray images, and the image-text contrastive pre-training further aligns the feature spaces, achieving better experimental results. Source code can be found on \url{https://github.com/Event-AHU/Medical_Image_Analysis}.

医学报告生成X光影像多阶段预训练CheXpert Plus

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。