首个外科影像大模型评测,揭示通用模型优于专用医疗模型。
Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study
- 构建首个大规模外科视觉语言模型评测体系。
- 通用模型在基础任务上表现接近常规场景,但需医学知识时性能下降。
- 专用医疗模型反而不如通用模型,提示其尚未适配复杂手术环境。
尽管传统计算机视觉模型在内窥镜领域长期表现不佳,但基础模型展现出良好的跨域泛化能力。本文首次系统评估了视觉语言模型(VLMs)在腹腔镜手术中的表现,涵盖多种先进模型、多个外科数据集及详尽的人工标注参考。研究聚焦三个核心问题:(1)当前VLM能否完成外科图像的基本感知任务?(2)能否处理基于帧的内窥镜场景理解任务?(3)专用医疗VLM与通用模型在此背景下孰优孰劣?结果表明,VLM可在物体计数与定位等基础任务中达到与通用领域相当的性能;但在依赖医学知识的任务中表现显著下滑。值得注意的是,专用医疗VLM在基本和高级任务中均落后于通用模型,说明其尚未充分适应手术环境的复杂性。本研究为下一代内窥镜AI系统的发展提供了关键洞见,并指明了医疗视觉语言模型亟待改进的方向。
原文摘要 · Abstract (English)
While traditional computer vision models have historically struggled to generalize to endoscopic domains, the emergence of foundation models has shown promising cross-domain performance. In this work, we present the first large-scale study assessing the capabilities of Vision Language Models (VLMs) for endoscopic tasks with a specific focus on laparoscopic surgery. Using a diverse set of state-of-the-art models, multiple surgical datasets, and extensive human reference annotations, we address three key research questions: (1) Can current VLMs solve basic perception tasks on surgical images? (2) Can they handle advanced frame-based endoscopic scene understanding tasks? and (3) How do specialized medical VLMs compare to generalist models in this context? Our results reveal that VLMs can effectively perform basic surgical perception tasks, such as object counting and localization, with performance levels comparable to general domain tasks. However, their performance deteriorates significantly when the tasks require medical knowledge. Notably, we find that specialized medical VLMs currently underperform compared to generalist models across both basic and advanced surgical tasks, suggesting that they are not yet optimized for the complexity of surgical environments. These findings highlight the need for further advancements to enable VLMs to handle the unique challenges posed by surgery. Overall, our work provides important insights for the development of next-generation endoscopic AI systems and identifies key areas for improvement in medical visual language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。