无需提示词的双变压器模型,自动分割手术场景
Surg-SegFormer: A Dual Transformer-Based Model for Holistic Surgical Scene Segmentation
- 采用双变压器架构,实现无需人工提示的全自动手术场景分割
- 在EndoVis2018上mIoU达0.80,在EndoVis2017上达0.54
- 适合手术教学与自动化分析,减轻专家指导负担
机器人辅助手术中的整体手术场景分割,使手术学员能够识别各类解剖组织、活动器械及关键结构(如静脉和血管)。由于术中时间紧迫,外科医生难以实时为学员提供详细操作场解释。这一挑战因专家外科医生数量远少于学员而加剧,导致明确划分可操作区与禁入区变得困难。因此,高性能语义分割模型可通过清晰的术后分析提供解决方案。然而,现有先进分割模型依赖用户生成的提示,不适用于常超过一小时的长视频手术录像。为此,我们提出Surg-SegFormer,一种新颖的免提示模型,其性能超越当前最先进方法。该模型在EndoVis2018数据集上达到0.80的平均交并比(mIoU),在EndoVis2017数据集上达到0.54。通过提供鲁棒且自动的手术场景理解,该模型显著减轻了专家外科医生的教学负担,使学员能独立有效地掌握复杂手术环境。
原文摘要 · Abstract (English)
Holistic surgical scene segmentation in robot-assisted surgery (RAS) enables surgical residents to identify various anatomical tissues, articulated tools, and critical structures, such as veins and vessels. Given the firm intraoperative time constraints, it is challenging for surgeons to provide detailed real-time explanations of the operative field for trainees. This challenge is compounded by the scarcity of expert surgeons relative to trainees, making the unambiguous delineation of go- and no-go zones inconvenient. Therefore, high-performance semantic segmentation models offer a solution by providing clear postoperative analyses of surgical procedures. However, recent advanced segmentation models rely on user-generated prompts, rendering them impractical for lengthy surgical videos that commonly exceed an hour. To address this challenge, we introduce Surg-SegFormer, a novel prompt-free model that outperforms current state-of-the-art techniques. Surg-SegFormer attained a mean Intersection over Union (mIoU) of 0.80 on the EndoVis2018 dataset and 0.54 on the EndoVis2017 dataset. By providing robust and automated surgical scene comprehension, this model significantly reduces the tutoring burden on expert surgeons, empowering residents to independently and effectively understand complex surgical environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。