轻量级语音转换让电子喉语音更自然、更易懂
Lightweight and perceptually-guided voice conversion for electro-laryngeal speech
- 移除音高能量模块,结合自监督预训练与有监督微调
- 字符错误率降低,自然度评分从1.1提升至3.3
- 适合语音康复研究者,尤其关注电子喉语音改善
电子喉(EL)语音具有恒定音高、有限语调和机械噪声,导致自然度和可懂度下降。本文提出对先进流式语音转换框架 StreamVC 的轻量级改进:移除音高与能量模块,结合自监督预训练与平行 EL 与健康语音(HE)数据的监督微调,并引入感知与可懂度损失进行引导。在不同损失配置下的客观与主观评估表明,最优模型(基于 WavLM 特征与人工反馈预测 +WavLM+HF)显著降低 EL 输入的字符错误率(CER),将自然度均值意见分(nMOS)从 1.1 提升至 3.3,且在所有评估指标上持续缩小与健康语音真值的差距。结果证明轻量级语音转换架构可用于 EL 语音康复,但语调生成与可懂度提升仍是主要挑战。
原文摘要 · Abstract (English)
Electro-laryngeal (EL) speech is characterized by constant pitch, limited prosody, and mechanical noise, reducing naturalness and intelligibility. We propose a lightweight adaptation of the state-of-the-art StreamVC framework to this setting by removing pitch and energy modules and combining self-supervised pretraining with supervised fine-tuning on parallel EL and healthy (HE) speech data, guided by perceptual and intelligibility losses. Objective and subjective evaluations across different loss configurations confirm their influence: the best model variant, based on WavLM features and human-feedback predictions (+WavLM+HF), drastically reduces character error rate (CER) of EL inputs, raises naturalness mean opinion score (nMOS) from 1.1 to 3.3, and consistently narrows the gap to HE ground-truth speech in all evaluated metrics. These findings demonstrate the feasibility of adapting lightweight voice conversion architectures to EL voice rehabilitation while also identifying prosody generation and intelligibility improvements as the main remaining bottlenecks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。