arXiv:2504.08274cs.SDcs.CL2025-04被引 4

统一模型实现多语言语音生成,细粒度控制发音风格。

Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation

  • 通过语言感知风格适配,统一处理多语言发音差异。
  • 支持跨语言细粒度声学风格控制,提升语音自然度。
  • 适合需要轻量级多语言语音合成的工业应用。

文本到语音(TTS)模型可将音素转换为波形,生成自然的人类语音。然而,多语言TTS仍面临音素词汇不一致及语调、说话风格差异的挑战。现有方法或为每种语言训练独立模型(性能高但资源消耗大),或使用统一模型(难以捕捉语言特异性风格)。本文提出LanStyleTTS,一种非自回归、语言感知风格自适应的TTS框架,标准化音素表示并实现跨语言音素级风格控制。该设计支持统一多语言模型,在无需训练语言专属模型的前提下生成准确高质量语音。我们将其集成至多种前沿非自回归TTS架构中进行评估,结果在不同模型主干上均表现出一致性能提升。此外,我们对比了梅尔频谱图与自编码器提取的潜在特征等声学表示形式。实验表明,潜在编码可显著降低模型规模与计算开销,同时保持高质量语音生成能力。

原文摘要 · Abstract (English)

Text-to-Speech (TTS) models can generate natural, human-like speech across multiple languages by transforming phonemes into waveforms. However, multilingual TTS remains challenging due to discrepancies in phoneme vocabularies and variations in prosody and speaking style across languages. Existing approaches either train separate models for each language, which achieve high performance at the cost of increased computational resources, or use a unified model for multiple languages that struggles to capture fine-grained, language-specific style variations. In this work, we propose LanStyleTTS, a non-autoregressive, language-aware style adaptive TTS framework that standardizes phoneme representations and enables fine-grained, phoneme-level style control across languages. This design supports a unified multilingual TTS model capable of producing accurate and high-quality speech without the need to train language-specific models. We evaluate LanStyleTTS by integrating it with several state-of-the-art non-autoregressive TTS architectures. Results show consistent performance improvements across different model backbones. Furthermore, we investigate a range of acoustic feature representations, including mel-spectrograms and autoencoder-derived latent features. Our experiments demonstrate that latent encodings can significantly reduce model size and computational cost while preserving high-quality speech generation.

语音合成多语言风格控制轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。