首个眼科手术动态三维重建数据集,助力精准分析手与器械交互。
Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery
- 构建多视角自动标注流程,融合运动先验与生物力学约束。
- 在710万帧数据上实现毫米级手部姿态与器械位姿重建。
- 适合医疗影像分析、微创手术机器人研究者参考。
精准的三维手部与器械重建对基于视觉的眼科微手术分析至关重要,但受限于缺乏真实且大规模的数据集和可靠的标注工具。本文提出OphNet-3D,首个面向眼科手术的大型RGB-D动态三维重建数据集,包含41段序列、40名外科医生,总计710万帧图像,提供12个手术阶段、10类器械、密集的MANO手部网格及完整的6自由度器械位姿标注。为实现高保真标注,设计多阶段自动标注流程,结合多视角观测、数据驱动运动先验、跨视图几何一致性、生物力学约束以及碰撞感知的交互约束。基于该数据集,建立双任务基准:双手姿态估计与手-器械交互重建,并提出H-Net与OH-Net两个专用模型。二者采用新型空间推理模块,结合弱透视相机建模与碰撞感知中心表示法,在手部重建上实现超过2mm的MPJPE提升,器械重建上达到最高23%的ADD-S改进。
原文摘要 · Abstract (English)
Accurate 3D reconstruction of hands and instruments is critical for vision-based analysis of ophthalmic microsurgery, yet progress has been hampered by the lack of realistic, large-scale datasets and reliable annotation tools. In this work, we introduce OphNet-3D, the first extensive RGB-D dynamic 3D reconstruction dataset for ophthalmic surgery, comprising 41 sequences from 40 surgeons and totaling 7.1 million frames, with fine-grained annotations of 12 surgical phases, 10 instrument categories, dense MANO hand meshes, and full 6-DoF instrument poses. To scalably produce high-fidelity labels, we design a multi-stage automatic annotation pipeline that integrates multi-view data observation, data-driven motion prior with cross-view geometric consistency and biomechanical constraints, along with a combination of collision-aware interaction constraints for instrument interactions. Building upon OphNet-3D, we establish two challenging benchmarks-bimanual hand pose estimation and hand-instrument interaction reconstruction-and propose two dedicated architectures: H-Net for dual-hand mesh recovery and OH-Net for joint reconstruction of two-hand-two-instrument interactions. These models leverage a novel spatial reasoning module with weak-perspective camera modeling and collision-aware center-based representation. Both architectures outperform existing methods by substantial margins, achieving improvements of over 2mm in Mean Per Joint Position Error (MPJPE) and up to 23% in ADD-S metrics for hand and instrument reconstruction, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。