arXiv:2508.19205cs.CLcs.AI2025-08被引 44

VibeVoice用扩散模型生成长达90分钟的多人对话,保持真实交流氛围。

VibeVoice Technical Report

  • 用扩散模型逐token生成语音隐向量,统一建模连续语音数据。
  • 自研语音分词器压缩率提升80倍,处理长语音效率大幅提高。
  • 适合需要高质量多说话人长对话生成的研究与应用者使用。

本报告介绍VibeVoice,一种通过下一词扩散(next-token diffusion)实现多说话人长时语音合成的新模型。该方法通过自回归生成潜在向量来统一建模连续数据。为此,我们提出一种新型连续语音分词器,相比流行的Encodec模型,数据压缩率提升80倍,同时保持相当的音质表现。该分词器在有效保留音频保真度的同时,显著提升长序列处理的计算效率。因此,VibeVoice可在64K上下文窗口长度下,支持最多4名说话人,生成长达90分钟的语音,精准捕捉真实对话氛围,性能超越开源及专有对话模型。

原文摘要 · Abstract (English)

This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeling continuous data by autoregressively generating latent vectors via diffusion. To enable this, we introduce a novel continuous speech tokenizer that, when compared to the popular Encodec model, improves data compression by 80 times while maintaining comparable performance. The tokenizer effectively preserves audio fidelity while significantly boosting computational efficiency for processing long sequences. Thus, VibeVoice can synthesize long-form speech for up to 90 minutes (in a 64K context window length) with a maximum of 4 speakers, capturing the authentic conversational ``vibe'' and surpassing open-source and proprietary dialogue models.

语音合成扩散模型多说话人长序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。