arXiv:2506.12570cs.SDcs.CL2025-06被引 9

StreamMel实现低延迟零样本语音合成,支持实时生成。

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling

  • 单阶段流式架构,交错文本与声学帧连续建模
  • 延迟更低,音质和说话人相似度优于现有流式系统
  • 适合实时语音大模型集成,无需离线处理

零样本语音合成近期取得显著进展,可在未见说话人上生成高质量语音,但多数系统因离线设计难以用于实时场景。现有流式方案多依赖多阶段流水线和离散表示,导致计算开销大、性能欠佳。本文提出StreamMel,首个单阶段流式语音合成框架,直接建模连续梅尔频谱图。通过交错文本标记与声学帧,实现低延迟自回归生成,同时保持高说话人相似度与自然度。在LibriSpeech上的实验表明,StreamMel在质量和延迟上均优于现有流式基线,性能接近离线系统,且支持高效实时生成,展现出与实时语音大模型集成的广阔前景。音频样例可访问:https://aka.ms/StreamMel。

原文摘要 · Abstract (English)

Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design. Current streaming TTS paradigms often rely on multi-stage pipelines and discrete representations, leading to increased computational cost and suboptimal system performance. In this work, we propose StreamMel, a pioneering single-stage streaming TTS framework that models continuous mel-spectrograms. By interleaving text tokens with acoustic frames, StreamMel enables low-latency, autoregressive synthesis while preserving high speaker similarity and naturalness. Experiments on LibriSpeech demonstrate that StreamMel outperforms existing streaming TTS baselines in both quality and latency. It even achieves performance comparable to offline systems while supporting efficient real-time generation, showcasing broad prospects for integration with real-time speech large language models. Audio samples are available at: https://aka.ms/StreamMel.

语音合成流式生成零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。