arXiv:2411.02591cs.CL2024-11被引 1

用表面肌电解码说话动作,揭示其几何结构

Geometry of orofacial neuromuscular signals: speech articulation decoding using surface electromyography

  • 将肌电信号映射到对称正定矩阵流形,用代数方法建模
  • 发现不同个体间肌电信号分布存在可量化差异
  • 适合资源受限设备,提升神经网络训练效率

本文提出一种基于表面肌电图(sEMG)解码语音发音的方法。研究采集面部、下颌和颈部的肌电信号,实现从肌电到语音的转换。结果表明,对称正定(SPD)矩阵流形是肌电信号的自然嵌入空间,可通过线性变换进行代数解释,并能分析跨个体的信号分布变化。该方法在数据和参数效率上表现优异,适用于肌电神经假体系统,尤其适合面临大规模数据收集困难且计算资源有限的嵌入式设备场景。

原文摘要 · Abstract (English)

Objective. In this article, we present data and methods for decoding speech articulations using surface electromyogram (EMG) signals. EMG-based speech neuroprostheses offer a promising approach for restoring audible speech in individuals who have lost the ability to speak intelligibly due to laryngectomy, neuromuscular diseases, stroke, or trauma-induced damage (e.g., from radiotherapy) to the speech articulators. Approach. To achieve this, we collect EMG signals from the face, jaw, and neck as subjects articulate speech, and we perform EMG-to-speech translation. Main results. Our findings reveal that the manifold of symmetric positive definite (SPD) matrices serves as a natural embedding space for EMG signals. Specifically, we provide an algebraic interpretation of the manifold-valued EMG data using linear transformations, and we analyze and quantify distribution shifts in EMG signals across individuals. Significance. Overall, our approach demonstrates significant potential for developing neural networks that are both data- and parameter-efficient, an important consideration for EMG-based systems, which face challenges in large-scale data collection and operate under limited computational resources on embedded devices.

肌电解码语音重建流形学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。