SARA通过双流变分自编码器融合语义与声学表征,提升零样本语音合成质量。
SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

- 双流架构分别处理语义锚点和残差声学信息,实现高效融合。
- 在零样本语音合成中生成更自然、更具表现力的语音,且推理加速下仍稳定。
- 无需复杂正则化,隐空间紧凑,适合实时应用或资源受限场景。
零样本文本到语音(TTS)依赖于鲁棒的语音表征。然而,当前语音分词器面临根本性权衡:声学编码器保留高保真音频但缺乏语言约束,导致生成时出现内容错误;而自监督学习(SSL)模型产生的语义标记虽能确保文本对齐精确,却丢失部分声学信息。为弥合这一差距,我们提出SARA,一种双流变分自编码器(VAE),直接融合冻结的SSL语义锚点与专用残差声学编码器。该方法有效缓解了上述困境,构建出高效紧凑的隐空间,且无需依赖复杂正则化。SARA在重建质量上超越多个强基线。此外,在下游零样本TTS任务中,其生成语音具有高度自然与表现力,并在加速推理下仍保持稳健性能,实现了合成速度与计算成本之间的良好权衡。
原文摘要 · Abstract (English)
Zero-shot text-to-speech (TTS) relies on robust speech representations. However, current speech tokenizers face a fundamental trade-off: acoustic codecs preserve high-fidelity audio but lack linguistic constraints, causing content errors during generation, whereas semantic tokens from self-supervised learning (SSL) models ensure precise text alignment but discard some acoustic information. To bridge this gap, we propose SARA, a dual-stream VAE that directly fuses a frozen SSL semantic anchor with a dedicated residual acoustic encoder. This effectively mitigates the dilemma, creating an efficient and compact latent space without relying on complex regularizers. SARA achieves superior reconstruction quality over strong baselines. Furthermore, in downstream zero-shot TTS tasks, it yields highly natural and expressive synthesis quality, and maintains robust generation performance even under accelerated inference, offering a favorable trade-off between synthesis speed and computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。