用表面肌电预测发音特征,还原可懂语音。
Articulatory Feature Prediction from Surface EMG during Speech Production
- 融合卷积与Transformer的模型,分步预测发音特征。
- 多数发音特征预测相关性达0.9,可生成可懂语音波形。
- 首个从表面肌电解码语音的方法,适合脑机接口研究者。
本文提出一种从说话过程中的表面肌电(sEMG)信号预测发音特征的模型。该模型结合卷积层与Transformer模块,并通过独立预测器输出发音特征。实验表明,该方法对大多数发音特征的预测相关性接近0.9。进一步验证了所预测的发音特征可解码为可理解的语音波形。据我们所知,这是首个通过发音特征从表面肌电解码语音波形的方法,为基于肌电的语音合成提供了新路径。此外,我们分析了电极位置与发音特征可预测性的关系,为优化肌电电极配置提供数据驱动的指导。代码与解码语音样本已公开。
原文摘要 · Abstract (English)
We present a model for predicting articulatory features from surface electromyography (EMG) signals during speech production. The proposed model integrates convolutional layers and a Transformer block, followed by separate predictors for articulatory features. Our approach achieves a high prediction correlation of approximately 0.9 for most articulatory features. Furthermore, we demonstrate that these predicted articulatory features can be decoded into intelligible speech waveforms. To our knowledge, this is the first method to decode speech waveforms from surface EMG via articulatory features, offering a novel approach to EMG-based speech synthesis. Additionally, we analyze the relationship between EMG electrode placement and articulatory feature predictability, providing knowledge-driven insights for optimizing EMG electrode configurations. The source code and decoded speech samples are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。