arXiv:2503.06984cs.SDcs.CV2025-03CVPR被引 10

用频谱分解技术让视频生成同步音频更精准

Synchronized Video-to-Audio Generation via Mel Quantization-Continuum Decomposition

  • 将频谱分解为三类信号,分别量化或连续处理
  • 在8项指标上达到当前最佳,音画同步更自然
  • 适合做音视频生成、可控音频合成的研究者

视频转音频生成对生成与视频同步的逼真音轨至关重要。本文提出一种新方法——梅尔频谱量化-连续分解(Mel-QCD),通过将梅尔频谱分解为三类不同信号,并对每类采用量化或连续处理,实现从视频中有效预测这些信号。设计的视频到全信号(V2X)预测器可准确捕捉视频中的音频驱动信息。预测结果经重构后,结合文本反演设计与ControlNet,用于控制成熟的文本到音频生成扩散模型。实验表明,该方法在8个评估指标上表现领先,涵盖音质、同步性及语义一致性等维度。代码与演示将公开于 https://wjc2830.github.io/MelQCD/。

原文摘要 · Abstract (English)

Video-to-audio generation is essential for synthesizing realistic audio tracks that synchronize effectively with silent videos. Following the perspective of extracting essential signals from videos that can precisely control the mature text-to-audio generative diffusion models, this paper presents how to balance the representation of mel-spectrograms in terms of completeness and complexity through a new approach called Mel Quantization-Continuum Decomposition (Mel-QCD). We decompose the mel-spectrogram into three distinct types of signals, employing quantization or continuity to them, we can effectively predict them from video by a devised video-to-all (V2X) predictor. Then, the predicted signals are recomposed and fed into a ControlNet, along with a textual inversion design, to control the audio generation process. Our proposed Mel-QCD method demonstrates state-of-the-art performance across eight metrics, evaluating dimensions such as quality, synchronization, and semantic consistency. Our codes and demos will be released at \href{Website}{https://wjc2830.github.io/MelQCD/}.

音视频生成扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。