arXiv:2503.08307cs.CV2025-03被引 4

提出新型Transformer架构,实现无限时长音视频生成与精准同步

$^R$FLAV: Rolling Flow matching for infinite Audio Video generation

  • 设计轻量级时序融合模块,高效对齐音视频模态
  • 在多模态生成任务中超越现有最先进模型性能
  • 适合需要高质量音视频协同生成的研究与应用

联合音视频(AV)生成仍是生成式AI中的重大挑战,主要源于三个关键要求:生成样本的质量、多模态无缝同步与时间连贯性,即音频需与视觉内容匹配,反之亦然,且支持无限视频时长。本文提出 $^R$-FLAV,一种新型基于Transformer的架构,全面应对上述挑战。我们探索了三种不同的跨模态交互模块,其中轻量级时序融合模块在对齐音视频模态方面表现最优且计算效率最高。实验结果表明,$^R$-FLAV 在多模态音视频生成任务中优于现有最先进模型。代码与检查点已公开于 https://github.com/ErgastiAlex/R-FLAV。

原文摘要 · Abstract (English)

Joint audio-video (AV) generation is still a significant challenge in generative AI, primarily due to three critical requirements: quality of the generated samples, seamless multimodal synchronization and temporal coherence, with audio tracks that match the visual data and vice versa, and limitless video duration. In this paper, we present $^R$-FLAV, a novel transformer-based architecture that addresses all the key challenges of AV generation. We explore three distinct cross modality interaction modules, with our lightweight temporal fusion module emerging as the most effective and computationally efficient approach for aligning audio and visual modalities. Our experimental results demonstrate that $^R$-FLAV outperforms existing state-of-the-art models in multimodal AV generation tasks. Our code and checkpoints are available at https://github.com/ErgastiAlex/R-FLAV.

音视频生成Transformer多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。