arXiv:2409.16431cs.CVcs.RO2024-09中稿 · IUS 2024

用3D CNN分析前臂超声视频,提升手势识别准确率

Hand Gesture Classification Based on Forearm Ultrasound Video Snippets Using 3D Convolutional Neural Networks

  • 用3D卷积捕捉超声视频的时空特征
  • 准确率达98.8%±0.9%,优于2D方法的96.5%±2.3%
  • 适合可穿戴交互与人机协同场景

基于超声的手部动作估计是人机交互中的关键研究方向。前臂超声能提供手部运动时肌肉形态变化的详细信息,可用于手势识别。以往工作多采用二维(2D)超声图像帧结合卷积神经网络(CNN)进行分析,但无法捕捉连续手部动作对应的时序特征。本研究采用3D卷积神经网络技术,提取超声视频片段中的时空模式以实现手势分类。我们对比了基于2D卷积、(2+1)D卷积、3D卷积及所提出网络的性能。结果表明,相比仅使用2D卷积层的网络(准确率96.5%±2.3%),所提方法将分类准确率提升至98.8%±0.9%。这些结果验证了使用超声视频片段在提升手势识别性能方面的优势。

原文摘要 · Abstract (English)

Ultrasound based hand movement estimation is a crucial area of research with applications in human-machine interaction. Forearm ultrasound offers detailed information about muscle morphology changes during hand movement which can be used to estimate hand gestures. Previous work has focused on analyzing 2-Dimensional (2D) ultrasound image frames using techniques such as convolutional neural networks (CNNs). However, such 2D techniques do not capture temporal features from segments of ultrasound data corresponding to continuous hand movements. This study uses 3D CNN based techniques to capture spatio-temporal patterns within ultrasound video segments for gesture recognition. We compared the performance of a 2D convolution-based network with (2+1)D convolution-based, 3D convolution-based, and our proposed network. Our methodology enhanced the gesture classification accuracy to 98.8 +/- 0.9%, from 96.5 +/- 2.3% compared to a network trained with 2D convolution layers. These results demonstrate the advantages of using ultrasound video snippets for improving hand gesture classification performance.

手势识别超声感知3D卷积人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。