arXiv:2605.00329cs.SDeess.AS2026-05被引 1

一拍生成高质量音频,速度比现有方法快8.5倍

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation

论文配图:Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation
图 1 · 摘自论文原文
  • 用能量评分头一步映射噪声到音频潜空间,跳过迭代采样
  • 通过蒸馏保留扩散模型的文本条件能力,保持生成质量
  • 适合需要低延迟音频生成的应用,如实时交互系统

自回归(AR)模型结合扩散头近期在文本到音频任务中表现优异,但其迭代解码和多步采样导致高延迟。为解决这一瓶颈,我们提出一种一步采样框架,结合能量-距离训练目标与表示级蒸馏。能量评分头将高斯噪声直接映射到音频潜变量,消除耗时的递归扩散采样过程;同时,从掩码自回归(MAR)文本到音频模型中蒸馏表示,保留扩散训练中学习到的强条件能力。在AudioCaps基准上,该方法在客观与主观指标上均持续优于现有一步基线(如ConsistencyTTA、SoundCTM、AudioLCM和AudioTurbo),显著缩小与多步采样自回归扩散系统的质量差距。相比最先进的AR扩散系统IMPACT,本方法实现最高达8.5倍的批量推理加速,同时保持高度竞争力的音频质量。结果表明,能量-距离训练与表示级蒸馏的结合,是实现快速、高质量文本到音频合成的有效方案。

原文摘要 · Abstract (English)

Autoregressive (AR) models with diffusion heads have recently achieved strong text-to-audio performance, yet their iterative decoding and multi-step sampling process introduce high-latency issues. To address this bottleneck, we propose a one-step sampling framework that combines an energy-distance training objective with representation-level distillation. An energy-scoring head maps Gaussian noise directly to audio latents in one step, eliminating the need for a costly recursive diffusion sampling process, while distillation from a masked autoregressive (MAR) text-to-audio model preserves the strong conditioning learned during diffusion training. On the AudioCaps benchmark, our method consistently outperforms prior one-step baselines such as ConsistencyTTA, SoundCTM, AudioLCM and AudioTurbo, on both objective and subjective metrics, while substantially narrowing the quality gap to AR diffusion systems with multi-step sampling. Compared to the state-of-the-art AR diffusion system, IMPACT, our approach achieves up to $8.5$x faster batch inference with highly competitive audio quality. These results demonstrate that combining energy-distance training with representation-level distillation provides an effective recipe for fast, high-quality text-to-audio synthesis.

音频生成扩散模型一拍采样高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。