用姿态数据增强肌电信号表征,提升手势识别泛化能力
CPEP: Contrastive Pose-EMG Pre-training Enhances Gesture Generalization on EMG Signals
- 通过对比学习对齐肌电与姿态特征,构建高质量肌电编码器
- 在分布内任务上性能领先基准21%,分布外任务提升72%
- 适合可穿戴设备上的零样本手势识别研究者
利用视频、图像和手部骨骼等高质量结构化数据进行手势分类是计算机视觉中的成熟课题。借助低功耗、低成本的生物信号(如表面肌电sEMG),可在可穿戴设备上实现连续手势预测。本文证明,从弱模态数据中学习与高质量结构化数据对齐的表征,能提升表征质量并支持零样本分类。我们提出对比姿态-肌电预训练(CPEP)框架,通过学习一个生成高质量、富含姿态信息的肌电表示的编码器,实现肌电与姿态表征的对齐。通过线性探测和零样本设置评估模型性能,结果表明,在分布内手势分类任务上,模型性能较emg2pose基准高出21%;在未见(分布外)手势分类任务上提升达72%。
原文摘要 · Abstract (English)
Hand gesture classification using high-quality structured data such as videos, images, and hand skeletons is a well-explored problem in computer vision. Leveraging low-power, cost-effective biosignals, e.g. surface electromyography (sEMG), allows for continuous gesture prediction on wearables. In this paper, we demonstrate that learning representations from weak-modality data that are aligned with those from structured, high-quality data can improve representation quality and enables zero-shot classification. Specifically, we propose a Contrastive Pose-EMG Pre-training (CPEP) framework to align EMG and pose representations, where we learn an EMG encoder that produces high-quality and pose-informative representations. We assess the gesture classification performance of our model through linear probing and zero-shot setups. Our model outperforms emg2pose benchmark models by up to 21% on in-distribution gesture classification and 72% on unseen (out-of-distribution) gesture classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。