arXiv:2601.15918cs.CV2026-01

无需微调的多视角手术手部姿态估计方法与首个大规模标注数据集。

A Multi-View Pipeline and Benchmark Dataset for 3D Hand Pose Estimation in Surgery

  • 基于多视角追踪与约束优化,仅用预训练模型实现鲁棒估计。
  • 2D关节误差降低31%,3D位置误差减少76%。
  • 适合手术动作分析、机器人辅助手术等临床研究场景。

目的:精准的3D手部姿态估计可支持手术技能评估、机器人辅助干预和几何感知工作流分析。然而,手术环境存在强烈且局部光照、器械或人员频繁遮挡、手套导致手部外观均一化等问题,加之缺乏标注数据,制约了可靠模型训练。方法:提出一种无需领域微调的多视角手术手部姿态估计管道,仅依赖现成预训练模型。该管道整合可靠的人体检测、全身姿态估计与跟踪区域上的先进2D手部关键点预测,再进行约束3D优化。此外,构建了一个新型手术基准数据集,包含超过68,000帧和3,000个手动标注的2D手部姿态,配有三角化生成的3D真值,录制于模拟手术室,涵盖不同场景复杂度。结果:定量实验表明,该方法持续优于基线,2D平均关节误差降低31%,3D每关节位置误差减少76%。结论:本工作为手术中的3D手部姿态估计建立了强基准,提供免训练管道与全面标注数据集,推动手术计算机视觉研究发展。

原文摘要 · Abstract (English)

Purpose: Accurate 3D hand pose estimation supports surgical applications such as skill assessment, robot-assisted interventions, and geometry-aware workflow analysis. However, surgical environments pose severe challenges, including intense and localized lighting, frequent occlusions by instruments or staff, and uniform hand appearance due to gloves, combined with a scarcity of annotated datasets for reliable model training. Method: We propose a robust multi-view pipeline for 3D hand pose estimation in surgical contexts that requires no domain-specific fine-tuning and relies solely on off-the-shelf pretrained models. The pipeline integrates reliable person detection, whole-body pose estimation, and state-of-the-art 2D hand keypoint prediction on tracked hand crops, followed by a constrained 3D optimization. In addition, we introduce a novel surgical benchmark dataset comprising over 68,000 frames and 3,000 manually annotated 2D hand poses with triangulated 3D ground truth, recorded in a replica operating room under varying levels of scene complexity. Results: Quantitative experiments demonstrate that our method consistently outperforms baselines, achieving a 31% reduction in 2D mean joint error and a 76% reduction in 3D mean per-joint position error. Conclusion: Our work establishes a strong baseline for 3D hand pose estimation in surgery, providing both a training-free pipeline and a comprehensive annotated dataset to facilitate future research in surgical computer vision.

3D姿态估计手术视觉多视角数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。