arXiv:2607.10421eess.AScs.SD2026-07

用流形约束提升单步语音生成质量,兼顾速度与保真度。

FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation

论文配图:FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation
图 1 · 摘自论文原文
  • 基于多表示弗雷歇距离损失,直接优化生成分布
  • 相比基线降低11.4%弗雷歇距离,FAD提升28.8%
  • 引入均值流锚定机制,解决后训练退化问题

尽管近期少步采样文本到音频生成模型如MeanAudio通过建模平均速度显著加速生成,其单步生成质量仍明显落后于多步模型。本文提出FdAudio,通过在预训练嵌入空间中利用多表示弗雷歇距离(FD)损失直接优化最终单步分布来缩小这一差距。为防止使用FD损失进行朴素后训练导致的多步退化问题,我们引入均值流一致性目标作为结构锚点。实验表明,FdAudio在少步系统中实现了领先的单步文本到音频生成质量,相较于基线MeanAudio框架,弗雷歇距离降低11.4%,FAD得分提升28.8%。特别地,通过提出均值流锚点,成功解决了FD后训练的多步退化问题,使25步采样路径能保持高保真音频合成,性能达到或超越强大多步模型,同时计算延迟仅为后者的极小部分。

原文摘要 · Abstract (English)

While recent few-step sampling text-to-audio generation models like MeanAudio substantially accelerate generation by modeling average velocities, their strict one-step generation quality still lags significantly behind multi-step counterparts. We propose FdAudio to bridge this gap. Unlike MeanAudio, which relies solely on regression against target velocity fields, our post-training approach optimizes the final one-step distribution directly across pre-trained embedding spaces via a multi-representation Fréchet-distance (FD) loss. Crucially, to prevent the multi-step degradation that naive post-training with FD-loss causes, we introduce a MeanFlow consistency objective as a structural anchor. Results demonstrate that FdAudio establishes state-of-the-art one-step T2A generation quality among few-step systems, yielding an 11.4% reduction in FD score and a 28.8% improvement in FAD score relative to the baseline MeanAudio framework. Notably, we solve FD post-training's naive multi-step degradation issue by proposing the MeanFlow anchor, enabling a 25-step sampling path to maintain high-fidelity audio synthesis that matches or surpasses strong multi-step models at a fraction of their computational latency.

文本转音频单步生成扩散模型音频合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。