用自然语言提示实现医学影像分割的个性化,让模型像专家一样灵活调整
ProSona: Prompt-Guided Personalization for Multi-Expert Medical Image Segmentation
- 通过潜空间建模专家标注风格,用提示词控制分割结果
- 在肺结节和前列腺MRI数据上,分割精度提升1%以上,误差降低17%
- 适合需要个性化医疗分析的研究者和临床医生使用
自动化医学图像分割面临高观察者间差异问题,尤其在肺结节勾画任务中专家意见常不一致。现有方法或融合为单一共识掩码,或为每位标注者设置独立模型分支。我们提出ProSona,一种两阶段框架,学习标注风格的连续潜空间,支持通过自然语言提示进行可控个性化。采用概率U-Net主干捕捉多样专家假设,结合提示引导投影机制在潜空间中生成个性化分割结果。多层级对比目标对齐文本与视觉表征,促进专家风格的解耦与可解释性。在LIDC-IDRI肺结节和多机构前列腺MRI数据集上,ProSona相比DPersona将广义能量距离降低17%,平均Dice提升超过1个百分点。结果表明,自然语言提示可提供灵活、准确且可解释的个性化分割控制。代码已公开。
原文摘要 · Abstract (English)
Automated medical image segmentation suffers from high inter-observer variability, particularly in tasks such as lung nodule delineation, where experts often disagree. Existing approaches either collapse this variability into a consensus mask or rely on separate model branches for each annotator. We introduce ProSona, a two-stage framework that learns a continuous latent space of annotation styles, enabling controllable personalization via natural language prompts. A probabilistic U-Net backbone captures diverse expert hypotheses, while a prompt-guided projection mechanism navigates this latent space to generate personalized segmentations. A multi-level contrastive objective aligns textual and visual representations, promoting disentangled and interpretable expert styles. Across the LIDC-IDRI lung nodule and multi-institutional prostate MRI datasets, ProSona reduces the Generalized Energy Distance by 17% and improves mean Dice by more than one point compared with DPersona. These results demonstrate that natural-language prompts can provide flexible, accurate, and interpretable control over personalized medical image segmentation. Our implementation is available online 1 .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。