arXiv:2509.14784eess.AS2025-09中稿 · ICASSP 2026被引 6

MELA-TTS用联合模型直接生成语音频谱,无需分步处理。

MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis

  • 融合Transformer与扩散模型,端到端生成语音频谱。
  • 引入对齐模块,提升文本与语音的语义一致性。
  • 支持零样本克隆,适合实时语音合成场景。

本文提出MELA-TTS,一种用于端到端文本到语音合成的新型联合Transformer-扩散框架。该模型通过自回归方式从语言和说话人条件生成连续的梅尔频谱帧,无需语音分词和多阶段处理流程。为解决连续特征建模的固有挑战,我们设计了表示对齐模块,在训练中将Transformer解码器的输出表示与预训练语音识别(ASR)编码器的语义嵌入进行对齐。该机制不仅加速了训练收敛,还增强了文本与声学域间的跨模态一致性。大量实验表明,MELA-TTS在多个评估指标上达到最先进水平,同时在离线与流式合成模式下均保持强大的零样本语音克隆能力。结果确立了连续特征生成方法在语音合成中的新基准,为基于离散符号的范式提供有力替代方案。

原文摘要 · Abstract (English)

This work introduces MELA-TTS, a novel joint transformer-diffusion framework for end-to-end text-to-speech synthesis. By autoregressively generating continuous mel-spectrogram frames from linguistic and speaker conditions, our architecture eliminates the need for speech tokenization and multi-stage processing pipelines. To address the inherent difficulties of modeling continuous features, we propose a representation alignment module that aligns output representations of the transformer decoder with semantic embeddings from a pretrained ASR encoder during training. This mechanism not only speeds up training convergence, but also enhances cross-modal coherence between the textual and acoustic domains. Comprehensive experiments demonstrate that MELA-TTS achieves state-of-the-art performance across multiple evaluation metrics while maintaining robust zero-shot voice cloning capabilities, in both offline and streaming synthesis modes. Our results establish a new benchmark for continuous feature generation approaches in TTS, offering a compelling alternative to discrete-token-based paradigms.

语音合成扩散模型端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。