arXiv:2510.11646cs.SD2025-10

用双表示法提升零样本语音合成速度与质量,减少生成步数同时保持音色自然。

BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis

  • 采用双语音表征机制,稀疏预测离散标记并重建连续特征
  • 合成速度显著提升,音质与说话人相似度达到领先水平
  • 适合追求高效高保真语音生成的研究与应用

自回归(AR)框架通过利用离散语音标记和大语言模型技术,在零样本文本到语音(TTS)合成中取得显著进展。然而现有 AR 型零样本 TTS 系统存在两大关键局限:(i) 固有的速度-质量权衡,顺序生成标记导致帧率降低或标记丰富性下降;(ii) 文本导向的监督不匹配,交叉熵损失对所有标记错误一视同仁,未考虑相邻标记间的细粒度声学相似性。为解决这些问题,我们提出 BridgeTTS,一种基于双语音表征范式 BridgeCode 的新型自回归语音合成框架。BridgeTTS 通过预测稀疏标记并重建丰富的连续特征,减少自回归迭代次数,同时联合优化标记级与特征级目标,进一步提升语音自然度与可懂性。实验表明,BridgeTTS 在保持竞争力音质和说话人相似度的同时,显著加速语音合成。语音演示见 https://test1562.github.io/demo/

原文摘要 · Abstract (English)

Autoregressive (AR) frameworks have recently achieved remarkable progress in zero-shot text-to-speech (TTS) by leveraging discrete speech tokens and large language model techniques. Despite their success, existing AR-based zero-shot TTS systems face two critical limitations: (i) an inherent speed-quality trade-off, as sequential token generation either reduces frame rates at the cost of expressiveness or enriches tokens at the cost of efficiency, and (ii) a text-oriented supervision mismatch, as cross-entropy loss penalizes token errors uniformly without considering the fine-grained acoustic similarity among adjacent tokens. To address these challenges, we propose BridgeTTS, a novel AR-TTS framework built upon the dual speech representation paradigm BridgeCode. BridgeTTS reduces AR iterations by predicting sparse tokens while reconstructing rich continuous features for high-quality synthesis. Joint optimization of token-level and feature-level objectives further enhances naturalness and intelligibility. Experiments demonstrate that BridgeTTS achieves competitive quality and speaker similarity while significantly accelerating synthesis. Speech demos are available at https://test1562.github.io/demo/.

语音合成自回归零样本双表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。