用几何不变量提升手术手势识别准确率
Multi-Modal Gesture Recognition from Video and Surgical Tool Pose Information via Motion Invariants
- 融合视频、工具位姿与曲率扭转度等几何不变量
- 在JIGSAWS数据集上达90.3%帧级准确率
- 适合做手术自动化与技能评估的研究者
实时识别手术手势是实现自动化手术活动识别、技能评估、术中辅助乃至手术自动化的关键。当前机器人手术系统提供了丰富的多模态数据,如视频与运动学信息。尽管近期多模态神经网络尝试学习视觉与运动学数据间的关系,但现有方法将运动学信息视为独立信号,忽略了工具末端位姿间的几何关联。然而,工具位姿具有几何相关性,其底层几何结构有助于神经网络学习手势表征。为此,我们提出将曲率和扭转度等运动不变量与视觉及运动学数据结合,通过关系图网络捕捉不同数据流间的内在联系。实验表明,加入不变量信号后,手势识别性能显著提升,在JIGSAWS缝合数据集上达到90.3%的帧级准确率。结果表明,结合位置与不变量的表示比传统的位置与四元数表示更优,凸显了在手势识别中对运动学进行几何感知建模的重要性。
原文摘要 · Abstract (English)
Recognizing surgical gestures in real-time is a stepping stone towards automated activity recognition, skill assessment, intra-operative assistance, and eventually surgical automation. The current robotic surgical systems provide us with rich multi-modal data such as video and kinematics. While some recent works in multi-modal neural networks learn the relationships between vision and kinematics data, current approaches treat kinematics information as independent signals, with no underlying relation between tool-tip poses. However, instrument poses are geometrically related, and the underlying geometry can aid neural networks in learning gesture representation. Therefore, we propose combining motion invariant measures (curvature and torsion) with vision and kinematics data using a relational graph network to capture the underlying relations between different data streams. We show that gesture recognition improves when combining invariant signals with tool position, achieving 90.3\% frame-wise accuracy on the JIGSAWS suturing dataset. Our results show that motion invariant signals coupled with position are better representations of gesture motion compared to traditional position and quaternion representations. Our results highlight the need for geometric-aware modeling of kinematics for gesture recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。