arXiv:2603.10541cs.CVcs.AI2026-03

评估11种基础模型在骨骼CT分割中对人类提示的敏感性,发现性能差异大且人类操作易出错。

Prompting with the human-touch: evaluating model-sensitivity of foundation models for musculoskeletal CT segmentation

  • 用2D/3D非迭代提示策略测试11个基础模型在4个解剖区域的表现
  • 2D最优为SAM和SAM2.1,3D最优为nnInteractive和Med-SAM2
  • 人类提示导致性能下降,提示变化影响大,适合临床医生参考

可提示基础模型(FMs)已革新医学图像分割。由于模型数量增多、评估数据集、指标和对比模型不一致,模型间直接比较困难,难以选择适合特定临床任务的模型。本研究在私有与公开数据集上,使用非迭代2D与3D提示策略,测试了11种可提示FM在腕部、肩部、髋部和下肢四个解剖区域的骨与植入物分割表现。通过专门的观察者研究收集人类提示,识别出帕累托最优模型并进一步分析。结果表明:1)不同模型与提示策略间分割性能差异显著;2)2D场景下最优为SAM与SAM2.1,3D场景下为nnInteractive与Med-SAM2;3)定位准确性和评分者一致性随解剖结构复杂度变化,简单结构(如腕骨)一致性高,复杂结构(如骨盆、胫骨、植入物)一致性低;4)使用人类提示时性能下降,说明基于参考标签提取的‘理想’提示可能夸大实际人机交互中的性能;5)所有模型均对提示变化敏感,尽管有两个模型表现出单人重复性鲁棒性,但无法推广至多人场景。结论:在人类驱动设置下选择最优基础模型仍具挑战性,即使高性能模型也对人类输入提示敏感。提示提取与模型推理代码已开源:https://github.com/CarolineMagg/segmentation-FM-benchmark/

原文摘要 · Abstract (English)

Promptable Foundation Models (FMs), initially introduced for natural image segmentation, have also revolutionized medical image segmentation. The increasing number of models, along with evaluations varying in datasets, metrics, and compared models, makes direct performance comparison between models difficult and complicates the selection of the most suitable model for specific clinical tasks. In our study, 11 promptable FMs are tested using non-iterative 2D and 3D prompting strategies on a private and public dataset focusing on bone and implant segmentation in four anatomical regions (wrist, shoulder, hip and lower leg). The Pareto-optimal models are identified and further analyzed using human prompts collected through a dedicated observer study. Our findings are: 1) The segmentation performance varies a lot between FMs and prompting strategies; 2) The Pareto-optimal models in 2D are SAM and SAM2.1, in 3D nnInteractive and Med-SAM2; 3) Localization accuracy and rater consistency vary with anatomical structures, with higher consistency for simple structures (wrist bones) and lower consistency for complex structures (pelvis, tibia, implants); 4) The segmentation performance drops using human prompts, suggesting that performance reported on "ideal" prompts extracted from reference labels might overestimate the performance in a human-driven setting; 5) All models were sensitive to prompt variations. While two models demonstrated intra-rater robustness, it did not scale to inter-rater settings. We conclude that the selection of the most optimal FM for a human-driven setting remains challenging, with even high-performing FMs being sensitive to variations in human input prompts. Our code base for prompt extraction and model inference is available: https://github.com/CarolineMagg/segmentation-FM-benchmark/

医学影像基础模型提示工程分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。