arXiv:2506.02555cs.CV2025-06被引 37

首个面向手术智能的多模态大模型,打通从视觉到推理的全流程理解。

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

  • 构建覆盖16类手术、7.79万次对话的大型手术数据库SurgVLM-DB
  • 在6个主流数据集上验证,大模型性能超越14款商用模型
  • 支持从感知到推理的全链条手术任务,适合医疗AI研究者使用

基础模型在生物医学领域取得突破性进展,但其在手术领域的应用仍不充分。手术智能需兼顾视觉感知、时间序列分析与高级推理,而现有通用多模态模型因缺乏领域专有监督和高质量数据支撑,难以满足需求。为此,我们提出SurgVLM——首个面向手术智能的大规模视觉语言基础模型,可统一处理多种手术任务。我们构建了包含181万帧图像与779万条对话的SurgVLM-DB,涵盖16类手术和18个解剖结构。整合23个公开数据集,统一标签并进行分层视觉-语言对齐,实现从视觉感知到高阶推理的渐进式覆盖。基于此,我们基于Qwen2.5-VL构建SurgVLM,经过10余项手术任务指令微调。进一步建立SurgVLM-Bench基准,包含6个常用手术数据集,覆盖多个关键下游任务。评估显示,SurgVLM(含7B、32B、72B三种版本)在该基准上表现优于14款主流商业模型(如GPT-4o、Gemini 2.0 Flash、Qwen2.5-Max)。

原文摘要 · Abstract (English)

Foundation models have achieved transformative success across biomedical domains by enabling holistic understanding of multimodal data. However, their application in surgery remains underexplored. Surgical intelligence presents unique challenges - requiring surgical visual perception, temporal analysis, and reasoning. Existing general-purpose vision-language models fail to address these needs due to insufficient domain-specific supervision and the lack of a large-scale high-quality surgical database. To bridge this gap, we propose SurgVLM, one of the first large vision-language foundation models for surgical intelligence, where this single universal model can tackle versatile surgical tasks. To enable this, we construct a large-scale multimodal surgical database, SurgVLM-DB, comprising over 1.81 million frames with 7.79 million conversations, spanning more than 16 surgical types and 18 anatomical structures. We unify and reorganize 23 public datasets across 10 surgical tasks, followed by standardizing labels and doing hierarchical vision-language alignment to facilitate comprehensive coverage of gradually finer-grained surgical tasks, from visual perception, temporal analysis, to high-level reasoning. Building upon this comprehensive dataset, we propose SurgVLM, which is built upon Qwen2.5-VL, and undergoes instruction tuning to 10+ surgical tasks. We further construct a surgical multimodal benchmark, SurgVLM-Bench, for method evaluation. SurgVLM-Bench consists of 6 popular and widely-used datasets in surgical domain, covering several crucial downstream tasks. Based on SurgVLM-Bench, we evaluate the performance of our SurgVLM (3 SurgVLM variants: SurgVLM-7B, SurgVLM-32B, and SurgVLM-72B), and conduct comprehensive comparisons with 14 mainstream commercial VLMs (e.g., GPT-4o, Gemini 2.0 Flash, Qwen2.5-Max).

手术智能多模态大模型视觉语言模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。