arXiv:2607.06405cs.MMcs.SD2026-07中稿 · ECCV

端到端生成同步音频,用新注意力机制提升音画对齐精度。

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space

论文配图:Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
图 1 · 摘自论文原文
  • 采用单阶段训练,通过改进注意力机制实现音画同步。
  • 在VGGSound上达到顶尖性能,零样本下音频质量超越闭源模型。
  • 自研数据增强管道生成声音导向描述,提升合成效果。

视频转音频(V2A)旨在生成与无声视频语义一致且时间同步的逼真音频。尽管近期取得进展,许多方法仍依赖多阶段训练,导致计算成本高、运行时间长,或需将视觉输入转换为文本以利用预训练文到音模型,牺牲了细粒度的时间线索。为此,我们提出Flowley,一种端到端、单阶段训练架构,通过结合视觉特征与文本提示生成配乐。关键在于引入渐进式软掩码交叉注意力机制,将音画同步嵌入注意力过程,相比标准注意力层无额外计算开销。我们还发现现有V2A基准缺乏声音导向的描述性标题,可能影响合成音频质量。为此,提出SoundCap——一个即插即用的管道,用于生成详尽的声音感知标题以指导模型。令人惊讶的是,不依赖任何预训练音视频对齐模块,Flowley在VGGSound多个指标上达到当前最优表现。进一步结合SoundCap,在零样本设置下,音频质量超越最强的现有闭源方法。

原文摘要 · Abstract (English)

Video-to-audio (V2A) generation aims to synthesize realistic audio that is both semantically consistent with and temporally synchronized to a silent video. Despite recent progress, many methods still rely on multi-stage training, resulting in high computational costs and long runtimes, or transform visual input into text to leverage pretrained text-to-audio models, sacrificing fine-grained temporal cues. To overcome these limitations, we propose Flowley, an end-to-end, single-stage training architecture that produces soundtracks by combining visual features with textual prompts. Crucially, we introduce Progressive Soft-masked Cross-Attention, which embeds audio-visual synchronization directly within its attention mechanism, adding zero additional computational cost compared to standard attention layers. We further observe that existing V2A benchmarks lack sound-oriented descriptive captions, which can potentially degrade the quality of the synthesized audio. To remedy this, we propose SoundCap, a plug-and-play pipeline for creating detailed, sound-aware captions that guide the model. Remarkably, without integrating any pretrained audio-visual alignment modules, Flowley achieves state-of-the-art performance on VGGSound across multiple metrics. Moreover, by incorporating SoundCap, we further exceed the performance of the strongest existing close-sourced methods in terms of audio quality in the zero-shot setting.

视频转音频音画对齐生成模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。