VoXtream2实现零样本端到端语音合成,支持实时语速调整。
VoXtream2: Full-stream TTS with dynamic speaking rate control
- 通过动态时长匹配与无分类器引导提升可控性
- 首包延迟仅74毫秒,全流模式速度达实时4倍
- 支持无需转录的文本无关音频提示,适合交互系统
面向交互式系统的零样本全流语音合成需在文本逐字到达时最小化延迟并保持可控性。我们提出VoXtream2,一种可在语句中实时更新的全流语音合成模型,具备动态语速控制能力。该模型结合时长状态上的分布匹配机制与跨条件信号的无分类器引导,提升合成质量与可控性。提示文本掩码技术实现无需转录的文本无关音频提示。在标准零样本基准和专用语速测试集上,尽管模型更小、训练数据更少,VoXtream2仍达到与公开基线相当的客观与主观性能。全流模式下,其运行速度达实时4倍,消费级显卡上首包延迟为74毫秒。
原文摘要 · Abstract (English)
Full-stream text-to-speech (TTS) for interactive systems must start speaking with minimal delay while remaining controllable as text arrives incrementally. We present VoXtream2, a zero-shot full-stream TTS model with dynamic speaking-rate control that can be updated mid-utterance on the fly. VoXtream2 combines a distribution matching mechanism over duration states with classifier-free guidance across conditioning signals to improve controllability and synthesis quality. Prompt-text masking enables textless audio prompting, removing the need for prompt transcription. Across standard zero-shot benchmarks and a dedicated speaking-rate test set, VoXtream2 achieves competitive objective and subjective results against public baselines despite a smaller model and less training data. In full-stream mode, it runs 4 times faster than real time with 74 ms first-packet latency on a consumer GPU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。