用自然语言自由描述器官,模型自动精准分割医学影像
Segment as You Wish -- Free-Form Language-Based Segmentation for Medical Images

- 基于领域语料生成真实临床描述的文本提示
- 支持解剖学、位置、大小等多类语言指令,精度超主流方法
- 能应对扫描方向变化,适合临床医生快速标注
医学影像对疾病诊断至关重要,准确分割有助于定位病灶。现有方法多依赖框或点提示,少有研究探索自然语言提示,而临床中医生常以语言描述观察与指令。为此,我们提出基于RAG的自由文本提示生成器,利用领域语料生成多样化真实描述;并引入FLanS模型,可处理包括解剖导向、位置驱动、大小驱动等各类自由文本提示。模型还包含对称性感知归一化模块,确保不同扫描方向下分割一致,减少解剖位置与图像外观混淆。FLanS在超过10万张来自7个公开数据集的医学图像上训练。实验表明,其在语言理解与分割精度上均优于现有最优方法,且在域内和跨域数据集上表现稳健。
原文摘要 · Abstract (English)
Medical imaging is crucial for diagnosing a patient's health condition, and accurate segmentation of these images is essential for isolating regions of interest to ensure precise diagnosis and treatment planning. Existing methods primarily rely on bounding boxes or point-based prompts, while few have explored text-related prompts, despite clinicians often describing their observations and instructions in natural language. To address this gap, we first propose a RAG-based free-form text prompt generator, that leverages the domain corpus to generate diverse and realistic descriptions. Then, we introduce FLanS, a novel medical image segmentation model that handles various free-form text prompts, including professional anatomy-informed queries, anatomy-agnostic position-driven queries, and anatomy-agnostic size-driven queries. Additionally, our model also incorporates a symmetry-aware canonicalization module to ensure consistent, accurate segmentations across varying scan orientations and reduce confusion between the anatomical position of an organ and its appearance in the scan. FLanS is trained on a large-scale dataset of over 100k medical images from 7 public datasets. Comprehensive experiments demonstrate the model's superior language understanding and segmentation precision, along with a deep comprehension of the relationship between them, outperforming SOTA baselines on both in-domain and out-of-domain datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。