arXiv:2504.02799cs.CVcs.AI2025-04被引 8

评估11个大模型在17项外科视觉任务中的表现,发现其泛化能力强但时空推理仍弱。

Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence

  • 用13个数据集测试11个先进多模态模型在手术视觉理解任务中的表现。
  • 上下文学习可使性能提升三倍,表明模型具备良好适应性。
  • 适合关注医疗AI通用能力与实际部署挑战的研究者参考。

大型视觉语言模型为人工智能驱动的图像理解提供了新范式,可在无需特定任务训练的情况下执行多种任务。这一灵活性在医学领域尤为突出,因为专家标注数据稀缺。然而,这些模型在以干预为导向的领域——尤其是外科手术——的实际应用价值尚不明确,因手术决策具有主观性且临床场景变化多样。本文对11个最先进的视觉语言模型在17个关键外科AI视觉理解任务(包括解剖识别、技能评估等)中进行了全面分析,覆盖腹腔镜、机器人和开放手术的13个数据集。实验表明,这些模型展现出良好的泛化能力,有时在非训练环境下表现优于监督学习模型。通过测试时引入示例(即上下文学习),性能最高提升三倍,说明适应能力是其核心优势。然而,需要空间或时间推理的任务仍难以处理。研究结果不仅对推动外科AI发展有启示,也为复杂动态临床及现实场景下的模型应用提供参考。

原文摘要 · Abstract (English)

Large Vision-Language Models offer a new paradigm for AI-driven image understanding, enabling models to perform tasks without task-specific training. This flexibility holds particular promise across medicine, where expert-annotated data is scarce. Yet, VLMs' practical utility in intervention-focused domains--especially surgery, where decision-making is subjective and clinical scenarios are variable--remains uncertain. Here, we present a comprehensive analysis of 11 state-of-the-art VLMs across 17 key visual understanding tasks in surgical AI--from anatomy recognition to skill assessment--using 13 datasets spanning laparoscopic, robotic, and open procedures. In our experiments, VLMs demonstrate promising generalizability, at times outperforming supervised models when deployed outside their training setting. In-context learning, incorporating examples during testing, boosted performance up to three-fold, suggesting adaptability as a key strength. Still, tasks requiring spatial or temporal reasoning remained difficult. Beyond surgery, our findings offer insights into VLMs' potential for tackling complex and dynamic scenarios in clinical and broader real-world applications.

视觉语言模型外科AI泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。