arXiv:2511.01345cs.CV2025-11

单点提示实现3D多病灶分割,提升医学影像标注效率

MIQ-SAM3D: From Single-Point Prompt to Multi-Instance Segmentation via Competitive Query Refinement

  • 用竞争性查询优化,将单点提示转为多实例查询
  • 在LiTS17和KiTS21上达可比性能,对提示鲁棒性强
  • 适合需要快速标注多个病灶的临床医学场景

精准的医学图像分割对肿瘤诊断与治疗规划至关重要。基于SAM的交互式分割因其强泛化能力受到关注,但多数方法遵循单点对单对象范式,限制了多病灶分割。此外,ViT主干网络虽能捕捉全局上下文,却常忽略高保真局部细节。我们提出MIQ-SAM3D,一种基于竞争性查询优化的多实例3D分割框架,实现从单点提示到多实例分割的转变。提示条件的实例查询生成器将单点提示转化为多个专用查询,使模型能从3D体积中检索出所有语义相似病灶。混合CNN-Transformer编码器通过空间门控将CNN提取的边界显著性注入ViT自注意力机制。竞争性优化的查询解码器则通过查询间竞争,实现端到端、并行的多实例预测。在LiTS17和KiTS21数据集上,MIQ-SAM3D达到可比性能,且对提示具有强鲁棒性,为临床相关多病灶病例的高效标注提供实用解决方案。

原文摘要 · Abstract (English)

Accurate segmentation of medical images is fundamental to tumor diagnosis and treatment planning. SAM-based interactive segmentation has gained attention for its strong generalization, but most methods follow a single-point-to-single-object paradigm, which limits multi-lesion segmentation. Moreover, ViT backbones capture global context but often miss high-fidelity local details. We propose MIQ-SAM3D, a multi-instance 3D segmentation framework with a competitive query optimization strategy that shifts from single-point-to-single-mask to single-point-to-multi-instance. A prompt-conditioned instance-query generator transforms a single point prompt into multiple specialized queries, enabling retrieval of all semantically similar lesions across the 3D volume from a single exemplar. A hybrid CNN-Transformer encoder injects CNN-derived boundary saliency into ViT self-attention via spatial gating. A competitively optimized query decoder then enables end-to-end, parallel, multi-instance prediction through inter-query competition. On LiTS17 and KiTS21 dataset, MIQ-SAM3D achieved comparable levels and exhibits strong robustness to prompts, providing a practical solution for efficient annotation of clinically relevant multi-lesion cases.

3D分割医学图像多实例SAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。