arXiv:2409.16415cs.CVcs.RO2024-09中稿 · IUS 2024被引 2

通过增量学习提升超声手势识别跨会话稳定性,准确率最高提升10%。

Improving Intersession Reproducibility for Forearm Ultrasound based Hand Gesture Classification through an Incremental Learning Approach

  • 用增量微调融合多会话数据,逐步优化分类模型
  • 经过两次微调后,分类准确率提升约10%
  • 适合开发可穿戴医疗设备的个性化人机交互系统

基于前臂超声图像的手势分类可用于人机交互系统。此前研究在单次会话中实现无探头移除的分类,但探头重新放置后性能下降,因分类器对探头位置敏感。本文提出利用多会话数据训练通用模型,通过增量微调实现持续优化。采集5种手势的超声数据,包括单会话内与跨会话数据。采用含5层级联卷积的卷积神经网络(CNN),预训练模型通过微调更新参数,卷积块作为特征提取器,其余层逐次增量更新。不同会话组合下的微调实验表明,增加微调会话数可提升准确率。每组实验经2次微调后,准确率平均提升约10%。结果证明,该方法可在降低存储、计算开销的同时提高准确率,适用于跨受试者泛化及个性化可穿戴设备开发。

原文摘要 · Abstract (English)

Ultrasound images of the forearm can be used to classify hand gestures towards developing human machine interfaces. In our previous work, we have demonstrated gesture classification using ultrasound on a single subject without removing the probe before evaluation. This has limitations in usage as once the probe is removed and replaced, the accuracy declines since the classifier performance is sensitive to the probe location on the arm. In this paper, we propose training a model on multiple data collection sessions to create a generalized model, utilizing incremental learning through fine tuning. Ultrasound data was acquired for 5 hand gestures within a session (without removing and putting the probe back on) and across sessions. A convolutional neural network (CNN) with 5 cascaded convolution layers was used for this study. A pre-trained CNN was fine tuned with the convolution blocks acting as a feature extractor, and the parameters of the remaining layers updated in an incremental fashion. Fine tuning was done using different session splits within a session and between multiple sessions. We found that incremental fine tuning can help enhance classification accuracy with more fine tuning sessions. After 2 fine tuning sessions for each experiment, we found an approximate 10% increase in classification accuracy. This work demonstrates that incremental learning through fine tuning on ultrasound based hand gesture classification can be used improves accuracy while saving storage, processing power, and time. It can be expanded to generalize between multiple subjects and towards developing personalized wearable devices.

手势识别超声成像增量学习可穿戴设备

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。