用视频帧间差异提升音效生成质量,无需额外模型。
Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation

- 以帧间时序差异作为核心视觉表征,增强音效生成条件。
- 在多个基准数据集上,生成音质超越依赖预训练的现有方法。
- 适合关注视频到音频生成、时序建模的研究者和开发者。
视频到音频(V2A)生成在图像到音频(I2A)基础上引入连续帧,提供关键时序线索以合成音频。然而,现有基于条件扩散的V2A方法通常依赖额外的音视频监督、声学结构预测或大模型推理,需添加额外网络或强先验。受视觉表征学习进展启发,我们提出TD-V2A,利用帧间时序差异(TD)作为区分V2A与I2A的核心表征,在最小修改架构的前提下丰富视觉条件。我们分别在帧级和特征级分析了TD的有效性,确定最优表示层级,并设计分层持续学习策略与渐进式时序差异引导方法,分别用于扩散训练与采样过程中的TD信息学习与利用。大量实验表明,通过本框架有效利用TD显著提升端到端V2A生成质量,甚至超越依赖对比音视频预训练等专用表示的方法。
原文摘要 · Abstract (English)
Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio synthesis. However, existing conditional diffusion-based V2A methods typically enhance visual conditioning with additional audio-visual supervision, acoustic structure prediction, or reasoning from large multimodal models, requiring extra networks or strong inductive biases. Inspired by recent advances in visual representation learning, we introduce TD-V2A, which leverages temporal differences (TD) as the key representation that distinguishes V2A from I2A, enriching visual conditioning with minimal architectural modification. We first investigate TD at both the frame and feature levels to identify the most effective representation level at which TD complements visual representations. Based on these findings, we develop a hierarchically continual learning strategy and an annealed temporal differences guidance method to progressively learn and exploit TD information during diffusion training and sampling process, respectively. Extensive experiments on benchmark datasets demonstrate that effectively exploiting TD through our proposed framework significantly improves end-to-end V2A generation quality, even outperforming dedicated V2A representations such as contrastive audio-visual pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。