用多帧上下文模型实现手术工具关键点高精度追踪
Video-based Surgical Tool-tip and Keypoint Tracking using Multi-frame Context-driven Deep Learning Models
- 基于多帧上下文的深度学习框架,提升关键点定位能力
- 在EndoVis数据集上达90%检测准确率,定位误差5.27像素
- 适用于手术技能评估与安全区划定,适合医疗AI研究者
机器人手术视频中自动化追踪手术工具关键点是技能评估、专家水平判断及安全区域划定等下游任务的关键。近年来,视觉领域的深度学习快速发展,推动了手术器械分割研究,但对工具尖端等特定关键点的追踪关注较少。本文提出一种新颖的多帧上下文驱动深度学习框架,用于定位和追踪手术视频中的工具关键点。模型在2015年EndoVis挑战赛标注帧上训练与测试,达到当前最优性能:关键点检测准确率达90%,定位均方根误差(RMS)为5.27像素。在自标注的更复杂场景的JIGSAWS数据集上,该方法对工具尖端和基底关键点的追踪误差整体低于4.2像素。该框架为手术器械关键点精准追踪提供了可行路径,可支持更多下游应用。项目与数据集网页:https://tinyurl.com/mfc-tracker
原文摘要 · Abstract (English)
Automated tracking of surgical tool keypoints in robotic surgery videos is an essential task for various downstream use cases such as skill assessment, expertise assessment, and the delineation of safety zones. In recent years, the explosion of deep learning for vision applications has led to many works in surgical instrument segmentation, while lesser focus has been on tracking specific tool keypoints, such as tool tips. In this work, we propose a novel, multi-frame context-driven deep learning framework to localize and track tool keypoints in surgical videos. We train and test our models on the annotated frames from the 2015 EndoVis Challenge dataset, resulting in state-of-the-art performance. By leveraging sophisticated deep learning models and multi-frame context, we achieve 90\% keypoint detection accuracy and a localization RMS error of 5.27 pixels. Results on a self-annotated JIGSAWS dataset with more challenging scenarios also show that the proposed multi-frame models can accurately track tool-tip and tool-base keypoints, with ${<}4.2$-pixel RMS error overall. Such a framework paves the way for accurately tracking surgical instrument keypoints, enabling further downstream use cases. Project and dataset webpage: https://tinyurl.com/mfc-tracker
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。