提出新方法评估医学影像分割对提示词的敏感性,提升模型可靠性判断能力。
Disentangling Prompt Dependence to Evaluate Segmentation Reliability in Gynecological MRI
- 首次解耦提示模糊性与局部敏感性,分离用户差异与交互误差
- 在两组妇科盆腔MRI数据上验证,两个指标与分割性能显著负相关
- 适合医疗领域需评估提示词影响的AI模型可靠性分析
可提示分割模型(如Segment Anything Models)实现了跨领域的通用零样本分割。尽管固定图像-提示对下的预测是确定性的,但这些模型对用户提示变化的鲁棒性(即提示依赖性)仍缺乏深入研究。在存在显著用户间差异的安全关键流程中,需要可解释、信息丰富的框架来评估提示依赖性。本文通过分析和测量提示变化对分割结果的影响,评估提示可提示分割的可靠性。我们提出了首个显式解耦提示模糊性(用户间差异)与局部敏感性(交互不精确)的提示依赖性定义,提供了可解释的分割鲁棒性视角。在两个用于子宫和膀胱分割的女性盆腔MRI数据集上的实验表明,这两个指标与分割性能呈强负相关,凸显了该框架在评估鲁棒性方面的价值。两项指标间互相关性低,支持了我们解耦设计的有效性,并为提示相关失败模式提供了有意义的指示。
原文摘要 · Abstract (English)
Promptable segmentation models (e.g., the Segment Anything Models) enable generalizable, zero-shot segmentation across diverse domains. Although predictions are deterministic for a fixed image-prompt pair, the robustness of these models to variations in user prompts, referred to as prompt dependence, remains underexplored. In safety-critical workflows with substantial inter-user variability, interpretable and informative frameworks are needed to evaluate prompt dependence. In this work, we assess the reliability of promptable segmentation by analyzing and measuring its sensitivity to prompt variability. We introduce the first formulation of prompt dependence that explicitly disentangles prompt ambiguity (inter-user variability) from local sensitivity (interaction imprecision), offering an interpretable view of segmentation robustness. Experiments on two female pelvic MRI datasets for uterus and bladder segmentation reveal a strong negative correlation between both metrics and segmentation performance, highlighting the value of our framework for assessing robustness. The two metrics have low mutual correlation, supporting the disentangled design of our formulation, and provide meaningful indicators of prompt-related failure modes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。