用专家诊断文本指导分割,提升唾液腺病变识别准确率
Multi-Sequence Parotid Gland Lesion Segmentation via Expert Text-Guided Segment Anything Model
- 用医生诊断文本自动生成提示信息,融合医学知识指导分割
- 在三个临床中心数据上达到当前最优效果,分割精度显著提升
- 适合需要结合医学经验的医疗影像分割场景
唾液腺病变分割对疾病治疗至关重要。由于病灶大小不一、边界复杂,精准分割仍具挑战。近期,基于Segment Anything Model(SAM)的微调在医学图像分割中表现优异,但其交互式分割严重依赖精确提示(如点、框、掩码),而真实临床中难以获取。此外,现有方法多为自动生成,忽视了医生的领域知识。为此,我们提出唾液腺分割任意模型(PG-SAM),通过专家诊断文本引导的提示生成模块,自动引入先验医学知识,指导分割过程;设计跨序列注意力模块,融合多模态互补信息以增强分割效果;将多序列图像特征与生成提示输入解码器,获得最终分割结果。实验表明,PG-SAM在三个独立临床中心的数据上均达到领先性能,验证了其临床适用性及诊断文本对真实场景下图像分割的有效性。
原文摘要 · Abstract (English)
Parotid gland lesion segmentation is essential for the treatment of parotid gland diseases. However, due to the variable size and complex lesion boundaries, accurate parotid gland lesion segmentation remains challenging. Recently, the Segment Anything Model (SAM) fine-tuning has shown remarkable performance in the field of medical image segmentation. Nevertheless, SAM's interaction segmentation model relies heavily on precise lesion prompts (points, boxes, masks, etc.), which are very difficult to obtain in real-world applications. Besides, current medical image segmentation methods are automatically generated, ignoring the domain knowledge of medical experts when performing segmentation. To address these limitations, we propose the parotid gland segment anything model (PG-SAM), an expert diagnosis text-guided SAM incorporating expert domain knowledge for cross-sequence parotid gland lesion segmentation. Specifically, we first propose an expert diagnosis report guided prompt generation module that can automatically generate prompt information containing the prior domain knowledge to guide the subsequent lesion segmentation process. Then, we introduce a cross-sequence attention module, which integrates the complementary information of different modalities to enhance the segmentation effect. Finally, the multi-sequence image features and generated prompts are feed into the decoder to get segmentation result. Experimental results demonstrate that PG-SAM achieves state-of-the-art performance in parotid gland lesion segmentation across three independent clinical centers, validating its clinical applicability and the effectiveness of diagnostic text for enhancing image segmentation in real-world clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。