通过压缩时域到25Hz潜空间,实现高效长对话语音合成。
ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching

- 将条件流匹配迁移到4倍时间压缩的潜空间中
- 峰值显存降低11.22倍,推理速度提升2.23倍
- 适合需要低内存、单次生成多分钟对话的应用
零样本对话语音合成虽受益于流匹配,但在密集梅尔频谱图上进行分钟级生成时会引发严重的内存瓶颈,常导致不自然的分段合成。本文提出ZipL-Dialog,将条件流匹配转移到4倍时间压缩(25 Hz)的潜空间中。为在压缩下保持声学保真度,采用具有辅助梅尔域监督的确定性梅尔自编码器,并优化ZipFormer的层级下采样策略。实验表明,与基线相比,ZipL-Dialog将最大峰值GPU显存降低11.22倍,推理速度提升2.23倍,显著降低单次生成多分钟对话的内存开销,同时保持感知自然性。
原文摘要 · Abstract (English)
Zero-shot dialog TTS benefits from flow-matching, but minute-scale generation on dense mel-spectrograms causes severe memory bottlenecks, often forcing unnatural chunked synthesis. We propose ZipL-Dialog, which shifts conditional flow-matching into a 4x time-compressed (25 Hz) latent space. To preserve acoustic fidelity under compression, we employ a deterministic mel autoencoder with auxiliary mel-domain supervision and optimize the ZipFormer's hierarchical downsampling schedule. Experiments show that ZipL-Dialog reduces maximum peak GPU memory by 11.22x and accelerates inference by 2.23x over the baseline, substantially lowering the memory footprint of single-pass multi-minute dialog synthesis while maintaining perceptual naturalness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。