多任务学习网络实现手术场景三维重建与器械识别
MT3DNet: Multi-Task learning Network for 3D Surgical Scene Reconstruction
- 设计多任务网络同步完成分割、深度估计和器械检测
- 在EndoVis2018数据集上实现高精度3D场景重建
- 适合需实时手术辅助的智能机器人系统研究者
在图像辅助的微创手术中,准确理解手术场景对实时反馈、技能评估及人机协作提升至关重要。本文提出一种新型多任务学习(MTL)网络,可同步完成高分辨率图像中手术场景的检测、分割、深度估计,并实现3D重建及器械分割与标注。为克服多任务优化难题,引入对抗性权重更新机制。该模型通过融合分割、深度估计与目标检测,显著提升手术场景理解能力,优于以往缺乏3D能力的研究。在EndoVis2018基准数据集上的全面实验验证了其在多项任务上的高效性与有效性。
原文摘要 · Abstract (English)
In image-assisted minimally invasive surgeries (MIS), understanding surgical scenes is vital for real-time feedback to surgeons, skill evaluation, and improving outcomes through collaborative human-robot procedures. Within this context, the challenge lies in accurately detecting, segmenting, and estimating the depth of surgical scenes depicted in high-resolution images, while simultaneously reconstructing the scene in 3D and providing segmentation of surgical instruments along with detection labels for each instrument. To address this challenge, a novel Multi-Task Learning (MTL) network is proposed for performing these tasks concurrently. A key aspect of this approach involves overcoming the optimization hurdles associated with handling multiple tasks concurrently by integrating a Adversarial Weight Update into the MTL framework, the proposed MTL model achieves 3D reconstruction through the integration of segmentation, depth estimation, and object detection, thereby enhancing the understanding of surgical scenes, which marks a significant advancement compared to existing studies that lack 3D capabilities. Comprehensive experiments on the EndoVis2018 benchmark dataset underscore the adeptness of the model in efficiently addressing all three tasks, demonstrating the efficacy of the proposed techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。