用Conformer模型提升无声语音接口的发音质量
Conformer-based Ultrasound-to-Speech Conversion
- 采用Conformer和带双向LSTM的Conformer结构处理超声波信号
- 听觉测试显示带LSTM的Conformer音质更优,基础版训练快3倍
- 适合追求音质提升或高效训练的语音接口研发者
深度神经网络在无声语音接口的超声波转语音任务中展现出潜力。本文采用两种基于Conformer的DNN架构(Base与带双向LSTM的版本),在Ultrasuite-Tal80数据集上对四位说话人进行个性化建模,生成的梅尔频谱通过HiFi-GAN声码器合成语音。相比标准2D-CNN基线,客观指标(MSE与梅尔倒谱失真)未显示统计显著提升。但MUSHRA听觉测试表明,带双向LSTM的Conformer感知质量更佳,而Conformer Base在性能持平基线的同时,因结构更简单实现3倍训练加速。结果表明,尤其是带双向LSTM的Conformer,是超声波转语音任务中替代CNN的有前景方案。
原文摘要 · Abstract (English)
Deep neural networks have shown promising potential for ultrasound-to-speech conversion task towards Silent Speech Interfaces. In this work, we applied two Conformer-based DNN architectures (Base and one with bi-LSTM) for this task. Speaker-specific models were trained on the data of four speakers from the Ultrasuite-Tal80 dataset, while the generated mel spectrograms were synthesized to audio waveform using a HiFi-GAN vocoder. Compared to a standard 2D-CNN baseline, objective measurements (MSE and mel cepstral distortion) showed no statistically significant improvement for either model. However, a MUSHRA listening test revealed that Conformer with bi-LSTM provided better perceptual quality, while Conformer Base matched the performance of the baseline along with a 3x faster training time due to its simpler architecture. These findings suggest that Conformer-based models, especially the Conformer with bi-LSTM, offer a promising alternative to CNNs for ultrasound-to-speech conversion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。