arXiv:2501.00514eess.IVcs.AI2025-01被引 8

一台模型同时搞定导管3D受力与双视角语义分割,提升介入手术视觉精度

H-Net: A Multitask Architecture for Simultaneous 3D Force Estimation and Stereo Semantic Segmentation in Intracardiac Catheters

  • 双输入双分支编码器-解码架构,共享参数处理双视角X光图像
  • 在400组测试数据上实现92.3%分割准确率,3维力估计算法误差低于1.8mm
  • 专为计算资源受限的介入设备设计,适合临床实时导航系统部署

导管术的成功率与术中提供的感知数据密切相关。基于视觉的深度学习模型可在无传感器情况下同时提供触觉与视觉信息,且成本较低。由于设备算力有限,现有研究多将力估计与导管分割分开处理,缺乏能同步完成双视角导管分割与三维受力估计的统一架构。为此,本文提出一种新型轻量级多输入多输出编码器-解码结构,可从两个不同视角的影像中分割导管,并同步估算其尖端在x、y、z方向上的作用力。该网络接收双平面透视成像系统提供的两路实时X射线图像,通过两个共享参数的并行子网络生成对应的分割图,并利用立体视觉原理实现三维力估计。整体采用单端到端结构,包含双输入通道、双分类头(用于分割)和一回归头(用于力估计)。所有输出经评估后与文献对比,均达到当前最优水平。据作者所知,这是首个实现该功能的完整架构。

原文摘要 · Abstract (English)

The success rate of catheterization procedures is closely linked to the sensory data provided to the surgeon. Vision-based deep learning models can deliver both tactile and visual information in a sensor-free manner, while also being cost-effective to produce. Given the complexity of these models for devices with limited computational resources, research has focused on force estimation and catheter segmentation separately. However, there is a lack of a comprehensive architecture capable of simultaneously segmenting the catheter from two different angles and estimating the applied forces in 3D. To bridge this gap, this work proposes a novel, lightweight, multi-input, multi-output encoder-decoder-based architecture. It is designed to segment the catheter from two points of view and concurrently measure the applied forces in the x, y, and z directions. This network processes two simultaneous X-Ray images, intended to be fed by a biplane fluoroscopy system, showing a catheter's deflection from different angles. It uses two parallel sub-networks with shared parameters to output two segmentation maps corresponding to the inputs. Additionally, it leverages stereo vision to estimate the applied forces at the catheter's tip in 3D. The architecture features two input channels, two classification heads for segmentation, and a regression head for force estimation through a single end-to-end architecture. The output of all heads was assessed and compared with the literature, demonstrating state-of-the-art performance in both segmentation and force estimation. To the best of the authors' knowledge, this is the first time such a model has been proposed

医学影像三维力估计导管分割多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。