用平均速度一步生成音频,快且保质。
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation
- 用平均速度代替逐帧流,实现一步生成
- 推理速度显著提升,音质与同步性不变
- 适合追求高效音视频生成的研究者
现有视频转音频方法在生成质量与推理效率间存在权衡。基于流匹配的模型依赖瞬时速度建模,需迭代采样,导致推理缓慢。为此,我们提出MeanFlow加速模型,通过平均速度表征流场,实现一步生成,显著提升多模态视频转音频(VTA)合成效率,同时保持音频质量、语义对齐与时间同步。此外,引入标量重缩放机制,在无分类器引导(CFG)时平衡条件与非条件预测,有效缓解一步生成中的畸变问题。由于音频生成网络与多模态条件联合训练,我们在文本转音频(TTA)任务上也进行了评估。实验表明,引入MeanFlow后,推理速度大幅提升,且在VTA与TTA任务上均未损失感知质量。
原文摘要 · Abstract (English)
A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow matching based models rely on modeling instantaneous velocity, inherently require an iterative sampling process, leading to slow inference speeds. To address this efficiency bottleneck, we introduce a MeanFlow-accelerated model that characterizes flow fields using average velocity, enabling one-step generation and thereby significantly accelerating multimodal video-to-audio (VTA) synthesis while preserving audio quality, semantic alignment, and temporal synchronization. Furthermore, a scalar rescaling mechanism is employed to balance conditional and unconditional predictions when classifier-free guidance (CFG) is applied, effectively mitigating CFG-induced distortions in one step generation. Since the audio synthesis network is jointly trained with multimodal conditions, we further evaluate it on text-to-audio (TTA) synthesis task. Experimental results demonstrate that incorporating MeanFlow into the network significantly improves inference speed without compromising perceptual quality on both VTA and TTA synthesis tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。