arXiv:2601.16895cs.CVcs.AI2026-01被引 1

评估大模型在手术器械检测中的表现,发现Qwen2.5效果最佳。

Evaluating Large Vision-language Models for Surgical Tool Detection

  • 对比三种大视觉语言模型在零样本和微调下的检测能力。
  • Qwen2.5在零样本和微调下均优于其他模型,识别率更高。
  • 适合关注手术AI、多模态模型应用的研究者阅读。

外科手术过程高度复杂,人工智能正成为支持手术引导与决策的变革力量。然而,当前多数AI系统为单模态,难以实现对手术流程的全面理解,亟需具备综合建模能力的通用外科AI系统。近年来,融合多模态数据处理的大视觉语言模型(VLM)展现出在手术任务建模与类人场景理解方面的潜力。尽管前景广阔,其在手术应用中的系统性研究仍有限。本研究评估了三种前沿VLM——Qwen2.5、LLaVA1.5与InternVL3.5——在GraSP机器人手术数据集上执行手术器械检测的效果,涵盖零样本与参数高效LoRA微调两种设置。结果表明,Qwen2.5在两种配置下均表现最优;相较于开集检测基线Grounding DINO,其在零样本泛化能力上更强,微调后性能相当。值得注意的是,Qwen2.5在器械识别方面更优,而Grounding DINO在定位精度上更具优势。

原文摘要 · Abstract (English)

Surgery is a highly complex process, and artificial intelligence has emerged as a transformative force in supporting surgical guidance and decision-making. However, the unimodal nature of most current AI systems limits their ability to achieve a holistic understanding of surgical workflows. This highlights the need for general-purpose surgical AI systems capable of comprehensively modeling the interrelated components of surgical scenes. Recent advances in large vision-language models that integrate multimodal data processing offer strong potential for modeling surgical tasks and providing human-like scene reasoning and understanding. Despite their promise, systematic investigations of VLMs in surgical applications remain limited. In this study, we evaluate the effectiveness of large VLMs for the fundamental surgical vision task of detecting surgical tools. Specifically, we investigate three state-of-the-art VLMs, Qwen2.5, LLaVA1.5, and InternVL3.5, on the GraSP robotic surgery dataset under both zero-shot and parameter-efficient LoRA fine-tuning settings. Our results demonstrate that Qwen2.5 consistently achieves superior detection performance in both configurations among the evaluated VLMs. Furthermore, compared with the open-set detection baseline Grounding DINO, Qwen2.5 exhibits stronger zero-shot generalization and comparable fine-tuned performance. Notably, Qwen2.5 shows superior instrument recognition, while Grounding DINO demonstrates stronger localization.

视觉语言模型手术检测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。