用用户涂鸦提示实现胸片多器官多病灶精准分割
PromptForSegCXR: Prompt-Driven Multi-Organ and Multi-Disease Segmentation in Chest X-rays using a Multi-stage Fusion Mechanism
- 输入胸片与用户涂鸦提示,通过多阶段融合提取特征
- 在23类病灶上达81.62%的Dice分数,优于现有模型10-23个百分点
- 轻量级设计适合临床交互使用,尤其适合标注效率要求高的场景
图像分割在自动化医学影像分析中至关重要,可精准识别解剖结构和病理区域。传统分割模型通常只针对单一器官或疾病,限制了其在临床中的适应性。尽管已有研究探索多器官、多疾病分割,但构建此类数据集需大量专家手动标注。基于提示的分割提供了灵活、用户引导的替代方案,可加速标注过程,但此前尚无工作解决胸片中跨多器官与多疾病、基于提示的交互式分割问题。本研究提出两项主要贡献:第一,构建了一个由专家设计的涂鸦提示数据集,涵盖23个类别(6个器官和17种疾病),数据源自多个公开胸片数据集;第二,提出PromptForSegCXR,一种轻量级双输入分割框架,将胸片与用户提供的涂鸦提示结合,准确分割多种解剖与病理区域。该模型采用多阶段特征融合策略整合空间与语义信息,并引入深度可分离点卷积与挤压-激励注意力模块,实现高效层次化特征提取与自适应重校准。实验结果表明,该模型在全数据集上达到81.62%的Dice分数,比基于SAM的提示分割模型最高提升10个百分点,比传统分割架构最高提升23个百分点,同时保持轻量化。这些结果验证了该方法在准确、灵活的提示驱动胸片分割中的有效性。
原文摘要 · Abstract (English)
Image segmentation is central to automated medical image analysis, enabling precise identification of anatomical structures and pathological regions. Conventional segmentation models typically target a single organ or disease, limiting their adaptability across clinical scenarios. While multi-organ and multi-disease segmentation has been explored, building such datasets requires extensive manual annotation by medical experts. Prompt-driven segmentation offers a flexible, user-guided alternative that speeds up annotation, yet no prior work has addressed prompt-based interactive segmentation across multiple organs and diseases in chest X-rays. This study makes two main contributions. First, we introduce a novel dataset of expert-designed doodle prompts spanning 23 classes (six organs and seventeen diseases), curated from multiple public chest X-ray datasets for prompt-driven segmentation. Second, we propose PromptForSegCXR, a lightweight dual-input segmentation framework that combines the chest X-ray with user-provided doodle prompts to accurately segment diverse anatomical and pathological regions. The model uses a multi-stage feature fusion strategy to integrate spatial and semantic representations, along with a depthwise-pointwise-residual convolution block with squeeze-and-excitation attention for efficient hierarchical feature extraction and adaptive recalibration. Experimental results show the model achieves a Dice score of 81.62 percent on the full dataset, outperforming SAM-based prompt segmentation models by up to 10 percent and conventional segmentation architectures by up to 23 percent, while remaining lightweight. These results demonstrate the effectiveness of the proposed approach for accurate, flexible, prompt-driven chest X-ray segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。