arXiv:2509.00866eess.IVcs.AI2025-09被引 1

通用模型在医学图像分割中表现如何?

Can General-Purpose Omnimodels Compete with Specialists? A Case Study in Medical Image Segmentation

  • 用通用模型与专用模型对比,测试零样本性能
  • 通用模型在难例上更稳定,易例上稍逊
  • 适合需要高鲁棒性的医疗场景使用

通用多模态模型能否在知识密集型领域与专用模型比肩?本研究在医疗图像分割领域开展对比实验,评估当前顶尖通用模型(Gemini,即“Nano Banana”模型)与特定领域深度学习模型在三种任务上的零样本表现:息肉(内窥镜)、视网膜血管(眼底)和乳腺肿瘤(超声)。研究通过专用模型的准确率筛选出“最易”和“最难”样本子集。结果显示:在息肉和乳腺肿瘤分割任务中,专用模型在易例上表现更优,但通用模型在难例上更具鲁棒性,专用模型常出现灾难性失效;而在细粒度的视网膜血管分割任务中,专用模型在所有样本上均保持领先。定性分析表明,通用模型可能具备更高敏感性,能识别人类标注者遗漏的细微解剖结构。结果表明,当前通用模型尚无法完全替代专用模型,但在处理复杂边缘案例时具有互补潜力。

原文摘要 · Abstract (English)

The emergence of powerful, general-purpose omnimodels capable of processing diverse data modalities has raised a critical question: can these ``jack-of-all-trades'' systems perform on par with highly specialized models in knowledge-intensive domains? This work investigates this question within the high-stakes field of medical image segmentation. We conduct a comparative study analyzing the zero-shot performance of a state-of-the-art omnimodel (Gemini, the ``Nano Banana'' model) against domain-specific deep learning models on three distinct tasks: polyp (endoscopy), retinal vessel (fundus), and breast tumor segmentation (ultrasound). Our study focuses on performance at the extremes by curating subsets of the ``easiest'' and ``hardest'' cases based on the specialist models' accuracy. Our findings reveal a nuanced and task-dependent landscape. For polyp and breast tumor segmentation, specialist models excel on easy samples, but the omnimodel demonstrates greater robustness on hard samples where specialists fail catastrophically. Conversely, for the fine-grained task of retinal vessel segmentation, the specialist model maintains superior performance across both easy and hard cases. Intriguingly, qualitative analysis suggests omnimodels may possess higher sensitivity, identifying subtle anatomical features missed by human annotators. Our results indicate that while current omnimodels are not yet a universal replacement for specialists, their unique strengths suggest a potential complementary role with specialist models, particularly in enhancing robustness on challenging edge cases.

医学图像通用模型分割鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。