arXiv:2502.16459eess.IVcs.AI2025-02综述被引 23

综述深度学习在手术视频分割与检测的进展,揭示大器官分割效果好、小结构仍难处理。

Deep learning approaches to surgical video segmentation and object detection: A Scoping Review

  • 系统梳理2014-2024年手术视频分割与检测的DL模型应用
  • 肝脏分割Dice达0.88,神经等小结构仅0.49,差距明显
  • 适用于实时处理(5-298帧/秒),但数据和泛化仍是瓶颈

计算机视觉在放射科、皮肤科和病理科已带来变革,但在手术应用中仍受限。本文综述2014至2024年间发表于PubMed、Embase和IEEE Xplore的58项研究,聚焦深度学习模型在手术视频中解剖结构的语义分割与目标检测性能。研究主要集中在普外科(34.4%)、结直肠手术(15.5%)和神经外科(13.8%),其中胆囊切除术(24.1%)和低位前切除术(8.6%)最常见。语义分割是主要任务(81%),常用模型为U-Net(24.1%)和DeepLab(22.4%)。大型器官如肝脏分割精度较高(Dice score: 0.88),而神经等小结构仅为0.49。多数模型具备实时推理能力,速度达5–298帧/秒。结论指出,尽管在大器官分割上取得显著进展并具备临床实时应用潜力,但小结构分割、数据可用性及模型泛化仍需突破。

原文摘要 · Abstract (English)

Introduction: Computer vision (CV) has had a transformative impact in biomedical fields such as radiology, dermatology, and pathology. Its real-world adoption in surgical applications, however, remains limited. We review the current state-of-the-art performance of deep learning (DL)-based CV models for segmentation and object detection of anatomical structures in videos obtained during surgical procedures. Methods: We conducted a scoping review of studies on semantic segmentation and object detection of anatomical structures published between 2014 and 2024 from 3 major databases - PubMed, Embase, and IEEE Xplore. The primary objective was to evaluate the state-of-the-art performance of semantic segmentation in surgical videos. Secondary objectives included examining DL models, progress toward clinical applications, and the specific challenges with segmentation of organs/tissues in surgical videos. Results: We identified 58 relevant published studies. These focused predominantly on procedures from general surgery [20(34.4%)], colorectal surgery [9(15.5%)], and neurosurgery [8(13.8%)]. Cholecystectomy [14(24.1%)] and low anterior rectal resection [5(8.6%)] were the most common procedures addressed. Semantic segmentation [47(81%)] was the primary CV task. U-Net [14(24.1%)] and DeepLab [13(22.4%)] were the most widely used models. Larger organs such as the liver (Dice score: 0.88) had higher accuracy compared to smaller structures such as nerves (Dice score: 0.49). Models demonstrated real-time inference potential ranging from 5-298 frames-per-second (fps). Conclusion: This review highlights the significant progress made in DL-based semantic segmentation for surgical videos with real-time applicability, particularly for larger organs. Addressing challenges with smaller structures, data availability, and generalizability remains crucial for future advancements.

手术视频语义分割深度学习医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。