arXiv:2502.14886cs.CV2025-02综述被引 13

综述基础模型在微创手术场景理解中的应用与挑战

Surgical Scene Understanding in the Era of Foundation AI Models: A Comprehensive Review

  • 整合CNN、ViT和SAM等模型提升手术分割与器械追踪精度
  • 基础模型显著改善手术阶段识别,但面临数据与算力挑战
  • 适合关注AI外科落地的临床研究者与医疗AI开发者

近年来,机器学习与深度学习的发展,特别是基础模型(FMs)的引入,显著提升了微创手术(MIS)中手术场景的理解能力。本文综述了卷积神经网络(CNN)、视觉变换器(ViTs)以及像分割任意模型(SAM)这样的基础模型在手术流程中的集成应用。这些技术提升了手术场景分割精度、器械跟踪和阶段识别能力。论文探讨了当前技术面临的挑战,如数据多样性、计算资源需求,以及临床环境中的伦理问题与集成障碍。通过强调基础模型的作用,本文将技术能力与临床需求相衔接,并提出未来研究方向,以增强AI在手术中的适应性、效率与伦理契合度。研究发现虽已取得显著进展,但仍需聚焦于实现技术在临床工作流中的无缝集成,从而提升手术精度、降低风险并优化患者预后。

原文摘要 · Abstract (English)

Recent advancements in machine learning (ML) and deep learning (DL), particularly through the introduction of Foundation Models (FMs), have significantly enhanced surgical scene understanding within minimally invasive surgery (MIS). This paper surveys the integration of state-of-the-art ML and DL technologies, including Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), and Foundation Models like the Segment Anything Model (SAM), into surgical workflows. These technologies improve segmentation accuracy, instrument tracking, and phase recognition in surgical scene understanding. The paper explores the challenges these technologies face, such as data variability and computational demands, and discusses ethical considerations and integration hurdles in clinical settings. Highlighting the roles of FMs, we bridge the technological capabilities with clinical needs and outline future research directions to enhance the adaptability, efficiency, and ethical alignment of AI applications in surgery. Our findings suggest that substantial progress has been made; however, more focused efforts are required to achieve seamless integration of these technologies into clinical workflows, ensuring they complement surgical practice by enhancing precision, reducing risks, and optimizing patient outcomes.

手术理解基础模型AI医疗计算机视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。