arXiv:2607.09134cs.SDcs.AI2026-07中稿 · ICML

ReGen通过分层多提示生成提升波形扩散模型效率与质量。

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

论文配图:ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models
图 1 · 摘自论文原文
  • 分层多提示框架联合建模表征与数据的向量场
  • 在12.5 Hz下实现高保真波形生成,6.25 Hz推理实时率达0.08
  • 适用于小样本语音合成,兼具高可懂度与说话人相似性

表示对齐(REPA)已被用于加速扩散模型训练,但我们发现,在扩散Transformer(DiT)中正则化中间表示可能隐式纠缠潜在变量,限制生成能力。为此,我们提出ReGen,一种分层多提示表示生成框架,可在单一扩散模型中联合估计表示与数据的多个向量场。我们进一步引入广义流匹配(GFM),以提升条件流匹配(CFM)的泛化能力。ReGen在单阶段波形扩散模型(包括神经音频编解码器和Wave-VAE)上验证有效,显著提升了在12.5 Hz压缩表示下的波形生成质量。我们还提出了ReGenVoice,一种基于潜扩散模型(LDM)的文语合成模型,在小数据集上实现了优异的语音可懂度(WER)和说话人相似性(SIM)。此外,该模型在6.25 Hz下运行,结合丰富的语义与声学潜表示,仅需4块GPU训练1天即可完成高效训练与快速推理,实时率(RTF)达0.08。音频样例见https://regenvoice.github.io/demo/。

原文摘要 · Abstract (English)

Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity. To address this issue, we propose ReGen, a hierarchical multi-prompt representation generation framework that jointly estimates multiple vector fields for both representations and data within a single diffusion model. We further introduce generalized flow matching (GFM) to improve the generalization of conditional flow matching (CFM). We validate ReGen on single-stage waveform diffusion models including neural audio codec and Wave-VAE. ReGen significantly improves waveform generation quality from highly compressed latent representations at 12.5 Hz. We also present ReGenVoice, a latent diffusion model (LDM)-based text-to-speech model that achieves strong speech intelligibility (WER) and speaker similarity (SIM) with a small dataset. Moreover, operating the LDM at 6.25 Hz with rich semantic and acoustic latent representation enables efficient training and sampling, requiring only 1 day of training on 4 GPUs and fast inference with an RTF of 0.08. Audio samples are available at https://regenvoice.github.io/demo/.

波形生成扩散模型语音合成高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。