arXiv:2511.11450cs.CVcs.LG2025-11被引 17

用自然语言描述自动分割3D医学影像,支持跨模态通用分割。

VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation

  • 通过多阶段视觉-语言融合,将自由文本映射为3D分割掩码。
  • 在未见数据集上零样本表现领先,对新类别有良好泛化能力。
  • 适合临床医生、研究人员快速实现个性化医学图像分析。

我们提出VoxTell,一种基于文本提示的三维医学图像分割视觉-语言模型。它能将从单个词到完整临床语句的自由描述转化为3D分割掩码。模型在超过62,000例CT、MRI和PET影像(涵盖上千种解剖与病理解剖类别)上训练,采用跨解码器层的多阶段视觉-语言融合,在多个尺度上对齐文本与视觉特征。在未见数据集上实现跨模态的最先进零样本性能,对熟悉概念表现优异,并能泛化至相关未见类别。大量实验表明其具备强跨模态迁移能力、对语言变化和临床表达的鲁棒性,以及从真实文本中精准生成实例级分割结果。代码已开源:https://www.github.com/MIC-DKFZ/VoxTell

原文摘要 · Abstract (English)

We introduce VoxTell, a vision-language model for text-prompted volumetric medical image segmentation. It maps free-form descriptions, from single words to full clinical sentences, to 3D masks. Trained on 62K+ CT, MRI, and PET volumes spanning over 1K anatomical and pathological classes, VoxTell uses multi-stage vision-language fusion across decoder layers to align textual and visual features at multiple scales. It achieves state-of-the-art zero-shot performance across modalities on unseen datasets, excelling on familiar concepts while generalizing to related unseen classes. Extensive experiments further demonstrate strong cross-modality transfer, robustness to linguistic variations and clinical language, as well as accurate instance-specific segmentation from real-world text. Code is available at: https://www.github.com/MIC-DKFZ/VoxTell

3D分割医学影像文本提示视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。