arXiv:2602.18959cs.CV2026-02

YOLOv10同时定位手术中双手并判断左右,提升紧急手术决策效率。

YOLOv10-Based Multi-Task Framework for Hand Localization and Laterality Classification in Surgical Videos

  • 基于YOLOv10设计多任务框架,联合检测手部位置与左右侧别。
  • 左/右手分类准确率分别为67%和71%,实时推理性能满足术中需求。
  • 适用于创伤外科视频分析,助力手-器械交互行为研究。

在创伤外科中,实时追踪手部对快速精准的术中决策至关重要。本文提出一种基于YOLOv10的多任务框架,可在复杂手术场景中同时定位双手并分类其左右侧(左或右)。模型在包含第一人称手术视频及标注手部边界框的Trauma THOMPSON Challenge 2025 Task 2数据集上训练,通过大量数据增强和多任务检测设计,提升了对运动模糊、光照变化及多样手部外观的鲁棒性。评估结果显示,左/右手分类准确率分别达67%和71%,而区分手部与背景仍具挑战;模型实现$ mAP_{[0.5:0.95]} $为0.33,并保持实时推理能力,展现出在术中部署的潜力。该工作为紧急手术中手-器械交互分析奠定了基础。

原文摘要 · Abstract (English)

Real-time hand tracking in trauma surgery is essential for supporting rapid and precise intraoperative decisions. We propose a YOLOv10-based framework that simultaneously localizes hands and classifies their laterality (left or right) in complex surgical scenes. The model is trained on the Trauma THOMPSON Challenge 2025 Task 2 dataset, consisting of first-person surgical videos with annotated hand bounding boxes. Extensive data augmentation and a multi-task detection design improve robustness against motion blur, lighting variations, and diverse hand appearances. Evaluation demonstrates accurate left-hand (67\%) and right-hand (71\%) classification, while distinguishing hands from the background remains challenging. The model achieves an $mAP_{[0.5:0.95]}$ of 0.33 and maintains real-time inference, highlighting its potential for intraoperative deployment. This work establishes a foundation for advanced hand-instrument interaction analysis in emergency surgical procedures.

手部定位多任务学习手术视觉YOLOv10

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。