arXiv:2603.27341cs.AIcs.CV2026-03

大模型在神经外科器械检测上仍表现不佳,规模扩展难见效。

A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling

  • 用2026年最先进视觉语言模型做神经外科器械检测
  • 多亿参数模型训练后性能提升有限,收益递减
  • 数据标注难、算力贵,非单纯扩规模能解决

近期人工智能模型在多项生物医学任务中已达到或超过人类专家水平,但手术类任务常被主流医疗基准套件忽略。由于手术需融合多种异质任务,通用型AI模型若能提升性能,或可成为理想协作工具。一方面,扩大模型规模和训练数据极具吸引力,每年产生数百万小时的手术视频;另一方面,手术数据准备需高专业度,训练成本高昂。这种权衡使现代AI在手术中的应用前景不明。本文通过2026年前沿AI方法在神经外科器械检测上的案例研究,发现即使使用多亿参数模型与大规模训练,当前视觉语言模型在看似简单的器械检测任务中仍表现不足。此外,扩展实验表明,增大模型规模和训练时间仅带来边际收益。结果暗示,现有模型在手术场景中仍面临重大障碍,且部分问题无法通过增加算力“解耦”,提示数据与标签并非唯一瓶颈。我们分析了主要制约因素,并提出潜在解决方案。

原文摘要 · Abstract (English)

Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but surgical benchmarks in particular are often missing from prominent medical benchmark suites. Since surgery requires integrating disparate tasks, generally-capable AI models could be particularly attractive as a collaborative tool if performance could be improved. On the one hand, the canonical approach of scaling architecture size and training data is attractive, especially since there are millions of hours of surgical video data generated per year. On the other hand, preparing surgical data for AI training requires significantly higher levels of professional expertise, and training on that data requires expensive computational resources. These trade-offs paint an uncertain picture of whether and to-what-extent modern AI could aid surgical practice. In this paper, we explore this question through a case study of surgical tool detection using state-of-the-art AI methods available in 2026. We demonstrate that even with multi-billion parameter models and extensive training, current Vision Language Models fall short in the seemingly simple task of tool detection in neurosurgery. Additionally, we show scaling experiments indicating that increasing model size and training time only leads to diminishing improvements in relevant performance metrics. Thus, our experiments suggest that current models could still face significant obstacles in surgical use cases. Moreover, some obstacles cannot be simply ``scaled away'' with additional compute and persist across diverse model architectures, raising the question of whether data and label availability are the only limiting factors. We discuss the main contributors to these constraints and advance potential solutions.

手术AI视觉语言模型器械检测规模效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。