arXiv:2510.23969cs.SDcs.CL2025-10ACL被引 2

用肌电图直接生成语音,无需复杂建模。

emg2speech: Synthesizing speech from electromyography using self-supervised speech models

  • 利用自监督语音模型提取肌电信号中的发音特征。
  • 线性映射使肌电功率预测相关性达0.85,可分离不同发音动作。
  • 适用于渐冻症患者无声发音转语音,端到端生成无须声码器训练。

我们提出一种神经肌肉语音接口,将说话时口面部肌群记录的肌电图(EMG)信号直接转换为音频。研究发现,自监督语音(S3)表示与肌肉电活动功率呈强线性关系:简单线性映射即可实现相关系数 r = 0.85 的功率预测。此外,不同发音动作对应的肌电功率向量形成结构化、可分离的聚类。这些结果表明,S3 模型隐式编码了发音机制,反映在肌电活动中。基于此结构,我们将肌电信号映射至 S3 表示空间并合成语音,实现端到端的肌电到语音生成,无需显式发音建模或声码器训练。我们在一名肌萎缩侧索硬化症(ALS)患者身上验证该系统,将其无声发音时采集的口面部肌电信号成功转化为语音。

原文摘要 · Abstract (English)

We present a neuromuscular speech interface that translates electromyographic (EMG) signals recorded from orofacial muscles during speech articulation directly into audio. We find that self-supervised speech (S3) representations are strongly linearly related to the electrical power of muscle activity: a simple linear mapping predicts EMG power from S3 representations with a correlation of r = 0.85. In addition, EMG power vectors associated with distinct articulatory gestures form structured, separable clusters. Together, these observations suggest that S3 models implicitly encode articulatory mechanisms, as reflected in EMG activity. Leveraging this structure, we map EMG signals into the S3 representation space and synthesize speech, enabling end-to-end EMG-to-speech generation without explicit articulatory modeling or vocoder training. We demonstrate this system with a participant with amyotrophic lateral sclerosis (ALS), converting orofacial EMG recorded while she silently articulated speech into audio.

肌电语音自监督模型神经接口语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。