用对齐特征实现音视频互转,统一框架更高效。
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

- 通过时序对齐自注意力融合音视频扩散模型特征
- 音视频同步效果显著优于更昂贵的基线模型
- 适合需要跨模态生成与同步的开发者
我们提出 AV-Link,一个统一的音视频生成框架,支持视频到音频(V2A)和音频到视频(A2V)双向生成。该框架利用冻结的视频与音频扩散模型激活特征,通过时序对齐的自注意力机制实现音视频间的双向信息交互。与以往分别使用专用模型且依赖预训练特征提取器的方法不同,AV-Link 在单一框架内直接利用互补模态的特征进行生成(如用视频特征生成音频,或用音频特征生成视频)。大量自动与主观评估表明,该方法在音视频同步性上取得显著提升,优于更复杂的基线模型如 MovieGen 视频到音频生成模型。
原文摘要 · Abstract (English)
We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioning. The key to our framework is a Fusion Block that facilitates bidirectional information exchange between video and audio diffusion models through temporally-aligned self attention operations. Unlike prior work that uses dedicated models for A2V and V2A tasks and relies on pretrained feature extractors, AV-Link achieves both tasks in a single self-contained framework, directly leveraging features obtained by the complementary modality (i.e. video features to generate audio, or audio features to generate video). Extensive automatic and subjective evaluations demonstrate that our method achieves a substantial improvement in audio-video synchronization, outperforming more expensive baselines such as the MovieGen video-to-audio model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。