用合成肌电数据提升语音重建准确率,解决真实数据少难题
Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
- 基于发音自信度筛选合成肌电信号,实现多说话人自训练
- 在测试中提升音素准确率,降低语音混淆和词错误率
- 适合关注神经接口与语音合成的研究者
有声肌电(V-ETS)模型从肌肉活动信号中重建语音,可用于神经喉科诊断等场景。然而,其发展受限于成对的肌电-语音数据稀缺。为此,我们提出一种基于置信度的多说话人自训练方法(CoM2S),并构建了新的人工数据集Libri-EMG。该方法利用预训练模型生成的合成肌电信号,通过基于音素级置信度的过滤机制,提升ETS模型性能。实验表明,该方法显著提高音素准确率,减少语音混淆,降低词错误率,验证了CoM2S的有效性。为支持未来研究,我们将开源代码及所提出的Libri-EMG数据集——一个公开、时序对齐、多说话人的有声肌电与语音记录数据集。
原文摘要 · Abstract (English)
Voiced Electromyography (EMG)-to-Speech (V-ETS) models reconstruct speech from muscle activity signals, facilitating applications such as neurolaryngologic diagnostics. Despite its potential, the advancement of V-ETS is hindered by a scarcity of paired EMG-speech data. To address this, we propose a novel Confidence-based Multi-Speaker Self-training (CoM2S) approach, along with a newly curated Libri-EMG dataset. This approach leverages synthetic EMG data generated by a pre-trained model, followed by a proposed filtering mechanism based on phoneme-level confidence to enhance the ETS model through the proposed self-training techniques. Experiments demonstrate our method improves phoneme accuracy, reduces phonological confusion, and lowers word error rate, confirming the effectiveness of our CoM2S approach for V-ETS. In support of future research, we will release the codes and the proposed Libri-EMG dataset-an open-access, time-aligned, multi-speaker voiced EMG and speech recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。