arXiv:2511.09592eess.IVq-bio.QM2025-11

基于不确定性感知的3D肿瘤分割模型,提升复杂医学影像中的精准识别能力。

Segment Any Tumour: An Uncertainty-Aware Vision Foundation Model for Whole-Body Analysis

  • 采用分层体积注意力与不确定性提示引导,增强模糊边界分割
  • 在11个数据集上表现优于现有模型,尤其在低对比度区域准确率显著提升
  • 支持交互式使用,适合临床医生快速辅助诊断

提示驱动的视觉基础模型(如Segment Anything Model)在计算机视觉中展现出强大适应性,但其直接应用于医学影像仍面临组织结构异质、成像伪影及低对比度边界等挑战,尤其是在肿瘤和癌症原发灶区域,导致模糊或重叠病灶分割效果不佳。本文提出轻量级三维体积基础模型SAT3D,旨在实现跨多种医学影像模态的鲁棒且可泛化的肿瘤分割。SAT3D结合移位窗口视觉变压器构建层次化体积表征,并引入不确定性感知训练流程,将不确定性估计作为提示,指导低对比度区域的可靠边界预测;对抗学习进一步提升了模型在模糊病理区域的表现。我们在11个公开数据集上进行评估,涵盖3,884例肿瘤及癌症病例用于训练,694例用于分布内测试。模型基于17,075对3D体数据-掩码样本,在多模态与多癌种下展现强泛化能力。为推动实际应用与临床转化,我们开发了3D Slicer插件,支持交互式提示驱动分割与可视化。大量实验表明,该模型在复杂及分布外场景下均有效提升分割精度,具备成为医学图像分析可扩展基础模型的潜力。

原文摘要 · Abstract (English)

Prompt-driven vision foundation models, such as the Segment Anything Model, have recently demonstrated remarkable adaptability in computer vision. However, their direct application to medical imaging remains challenging due to heterogeneous tissue structures, imaging artefacts, and low-contrast boundaries, particularly in tumours and cancer primaries leading to suboptimal segmentation in ambiguous or overlapping lesion regions. Here, we present Segment Any Tumour 3D (SAT3D), a lightweight volumetric foundation model designed to enable robust and generalisable tumour segmentation across diverse medical imaging modalities. SAT3D integrates a shifted-window vision transformer for hierarchical volumetric representation with an uncertainty-aware training pipeline that explicitly incorporates uncertainty estimates as prompts to guide reliable boundary prediction in low-contrast regions. Adversarial learning further enhances model performance for the ambiguous pathological regions. We benchmark SAT3D against three recent vision foundation models and nnUNet across 11 publicly available datasets, encompassing 3,884 tumour and cancer cases for training and 694 cases for in-distribution evaluation. Trained on 17,075 3D volume-mask pairs across multiple modalities and cancer primaries, SAT3D demonstrates strong generalisation and robustness. To facilitate practical use and clinical translation, we developed a 3D Slicer plugin that enables interactive, prompt-driven segmentation and visualisation using the trained SAT3D model. Extensive experiments highlight its effectiveness in improving segmentation accuracy under challenging and out-of-distribution scenarios, underscoring its potential as a scalable foundation model for medical image analysis.

肿瘤分割3D模型不确定性感知医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。