arXiv:2505.07701cs.SDcs.AI2025-05被引 8

轻量级端到端语音合成,低资源设备也能实时生成自然语音。

Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications

  • 设计轻量级端到端模型,直接从文本生成波形。
  • 参数量减少90%,实时因子提升10倍,性能达顶尖水平。
  • 适合移动端低资源场景,且优于传统两阶段训练方式。

近期研究显示,以端到端(E2E)方式直接从文本建模原始波形,生成的语音比基于级联或两阶段的传统神经语音合成系统更自然。然而,当前最先进的端到端模型计算复杂且内存占用高,难以在低资源环境下实现实时离线设备应用。为此,我们提出轻量级端到端语音合成(LE2E)模型,可在极低计算资源下生成高质量语音。我们在LJSpeech数据集上评估该模型,结果表明其性能达到最先进水平,参数量最多减少90%,实时因子提升10倍。此外,我们证明了端到端训练范式相比等效两阶段架构在语音质量上更具优势。实验结果表明,LE2E是面向设备端实时、高质量、低资源语音合成应用的有力方案。

原文摘要 · Abstract (English)

Recent works have shown that modelling raw waveform directly from text in an end-to-end (E2E) fashion produces more natural-sounding speech than traditional neural text-to-speech (TTS) systems based on a cascade or two-stage approach. However, current E2E state-of-the-art models are computationally complex and memory-consuming, making them unsuitable for real-time offline on-device applications in low-resource scenarios. To address this issue, we propose a Lightweight E2E-TTS (LE2E) model that generates high-quality speech requiring minimal computational resources. We evaluate the proposed model on the LJSpeech dataset and show that it achieves state-of-the-art performance while being up to $90\%$ smaller in terms of model parameters and $10\times$ faster in real-time-factor. Furthermore, we demonstrate that the proposed E2E training paradigm achieves better quality compared to an equivalent architecture trained in a two-stage approach. Our results suggest that LE2E is a promising approach for developing real-time, high quality, low-resource TTS applications for on-device applications.

语音合成轻量模型端到端低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。