用分割掩码提升医学影像报告生成的细粒度理解能力
MAIRA-Seg: Enhancing Radiology Report Generation with Segmentation-Aware Multimodal Large Language Models
- 引入可训练的分割标记提取器,融合语义分割掩码增强多模态模型感知
- 在MIMIC-CXR数据集上优于无分割基线,报告生成质量显著提升
- 适合关注医学影像分析与临床辅助决策的研究者与开发者
近年来,人工智能在放射科报告生成领域受到广泛关注,尤其针对胸部X光片(CXRs)。本文探究将像素级信息通过分割掩码引入多模态大语言模型(MLLMs),是否能提升其对医学图像的细粒度解读能力。为此,我们提出MAIRA-Seg框架,利用语义分割掩码与胸部X光片联合生成放射科报告。首先训练专家分割模型,获取胸部影像中特定解剖结构的伪标签掩码;随后基于专用于胸部X光报告生成的MAIRA模型架构,集成一个可训练的分割标记提取模块,并采用掩码感知提示策略生成初步报告。在公开可用的MIMIC-CXR数据集上的实验表明,MAIRA-Seg优于不使用分割掩码的基线模型。我们还研究了多掩码提示策略,结果发现MAIRA-Seg始终表现相当或更优。结果证实,引入分割掩码有助于增强MLLM的细腻推理能力,可能带来更优的临床效果。
原文摘要 · Abstract (English)
There is growing interest in applying AI to radiology report generation, particularly for chest X-rays (CXRs). This paper investigates whether incorporating pixel-level information through segmentation masks can improve fine-grained image interpretation of multimodal large language models (MLLMs) for radiology report generation. We introduce MAIRA-Seg, a segmentation-aware MLLM framework designed to utilize semantic segmentation masks alongside CXRs for generating radiology reports. We train expert segmentation models to obtain mask pseudolabels for radiology-specific structures in CXRs. Subsequently, building on the architectures of MAIRA, a CXR-specialised model for report generation, we integrate a trainable segmentation tokens extractor that leverages these mask pseudolabels, and employ mask-aware prompting to generate draft radiology reports. Our experiments on the publicly available MIMIC-CXR dataset show that MAIRA-Seg outperforms non-segmentation baselines. We also investigate set-of-marks prompting with MAIRA and find that MAIRA-Seg consistently demonstrates comparable or superior performance. The results confirm that using segmentation masks enhances the nuanced reasoning of MLLMs, potentially contributing to better clinical outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。