让肌电信号能用自然语言查询,实现手部动作的双向检索。
MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding

- 构建多视角动作语义空间,将肌电信号映射到统一语义空间。
- 在多个数据集上实现优于基线的双向检索效果,跨用户泛化能力强。
- 适合需要自然语言控制假肢或可穿戴设备的研究者使用。
肌电图(EMG)直接反映肌肉激活,是手势识别、假肢控制和可穿戴交互的关键传感模态。现有方法通常将手部动作理解为固定标签的分类任务,难以支持基于动作描述的查询、检索和泛化。本文提出 MyoSem,一个将肌电信号与自然语言动作语义对齐的框架,通过多视角动作语义构建、激活感知的肌电编码和语义查询对齐,实现肌电信号与文本描述间的双向检索。我们在 EMG2Pose 和 NinaPro 系列数据集上系统评估了 MyoSem。结果表明,该方法在肌电-文本双向检索中表现优异,普遍优于多数基线模型,并在未见用户、未见动作类别及截肢用户迁移场景中展现出良好泛化能力。消融实验与可视化进一步验证了各模块的有效性。总体而言,MyoSem 将基于肌电的手部动作理解从固定标签识别推进至可查询的双向语义检索,为语言驱动的肌电动作理解提供了新范式。
原文摘要 · Abstract (English)
Electromyography (EMG) directly reflects muscle activation and is a key sensing modality for gesture recognition, prosthetic control, and wearable interaction. Existing EMG methods, however, commonly formulate hand action understanding as classification over fixed labels, making it difficult to support querying, retrieval, and generalization based on action descriptions. We present MyoSem, an EMG--action semantic alignment framework that maps low-level EMG signals into a shared semantic space constructed from multi-view action descriptions. MyoSem combines multi-view action-semantic construction, activation-aware EMG encoding, and semantic query alignment, enabling bidirectional retrieval between EMG signals and text descriptions. We systematically evaluate MyoSem on EMG2Pose and NinaPro-series datasets. Results show that MyoSem performs well on EMG--text bidirectional retrieval, generally outperforms most baselines, and shows favorable generalization to unseen users, held-out action classes, and amputee-user transfer scenarios. Ablations and visualizations further validate the effectiveness of each module. Overall, MyoSem advances EMG-based hand action understanding from fixed-label recognition toward queryable bidirectional semantic retrieval, providing a new modeling paradigm for language-mediated EMG action understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。