对比八种模型在机器人手术中的解剖结构识别表现,发现注意力模型更优。
Benchmarking Pretrained Attention-based Models for Real-Time Recognition in Robot-Assisted Esophagectomy
- 用注意力机制建模手术中复杂结构的全局关系。
- 预训练ADE20k比ImageNet效果更好,尤其提升小类识别。
- Mask2Former在分割精度和距离度量上均领先,适合新手导航辅助。
食管癌是全球最常见的癌症之一,传统治疗为开腹食管切除术,近年机器人辅助微创食管切除术(RAMIE)成为有前景的替代方案。然而对新手外科医生而言,手术中易出现空间定向障碍。计算机辅助解剖识别有望改善手术导航,但相关研究仍有限。本研究构建了迄今最大规模的RAMIE语义分割数据集,涵盖最多的重要解剖结构与手术器械。该任务面临类别不平衡及神经等复杂结构识别难题。我们评估了八种实时深度学习模型,使用两个预训练数据集进行对比。涵盖传统与注意力网络,假设注意力模型能更好捕捉全局模式并应对血液或组织遮挡。基准测试包含自建的RAMIE数据集与公开的CholecSeg8k数据集,全面评估手术分割性能。结果表明,基于ADE20k预训练优于ImageNet,注意力模型整体表现更优,其中SegNeXt与Mask2Former获得更高Dice分数,而Mask2Former在平均对称表面距离指标上也更优。
原文摘要 · Abstract (English)
Esophageal cancer is among the most common types of cancer worldwide. It is traditionally treated using open esophagectomy, but in recent years, robot-assisted minimally invasive esophagectomy (RAMIE) has emerged as a promising alternative. However, robot-assisted surgery can be challenging for novice surgeons, as they often suffer from a loss of spatial orientation. Computer-aided anatomy recognition holds promise for improving surgical navigation, but research in this area remains limited. In this study, we developed a comprehensive dataset for semantic segmentation in RAMIE, featuring the largest collection of vital anatomical structures and surgical instruments to date. Handling this diverse set of classes presents challenges, including class imbalance and the recognition of complex structures such as nerves. This study aims to understand the challenges and limitations of current state-of-the-art algorithms on this novel dataset and problem. Therefore, we benchmarked eight real-time deep learning models using two pretraining datasets. We assessed both traditional and attention-based networks, hypothesizing that attention-based networks better capture global patterns and address challenges such as occlusion caused by blood or other tissues. The benchmark includes our RAMIE dataset and the publicly available CholecSeg8k dataset, enabling a thorough assessment of surgical segmentation tasks. Our findings indicate that pretraining on ADE20k, a dataset for semantic segmentation, is more effective than pretraining on ImageNet. Furthermore, attention-based models outperform traditional convolutional neural networks, with SegNeXt and Mask2Former achieving higher Dice scores, and Mask2Former additionally excelling in average symmetric surface distance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。