arXiv:2409.08601cs.SDcs.MM2024-09被引 25

让视频生成更真实声音,精准对齐语义与时间。

STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

  • 用时序与语义双特征引导音频生成。
  • 生成音频在质量、语义一致性和时间对齐上超越现有模型。
  • 适合音视频合成、虚拟场景制作等应用。

视觉与听觉是人类感知世界的重要方式。尽管文本到视频生成在过去一年取得显著进展,但生成视频缺乏协调的音频,限制了其广泛应用。本文提出语义与时间对齐的视频转音频方法(STA-V2A),通过提取视频的局部时序特征和全局语义特征,并结合文本作为跨模态引导,提升音频生成效果。为解决视频信息冗余问题,设计了启始点预测预训练任务以提取局部时序特征,以及注意力池化模块以提取全局语义特征。为弥补视频中语义信息不足,引入基于文本到音频先验初始化的潜在扩散模型及跨模态引导机制。此外,提出新的音频-音频对齐度量指标 Audio-Audio Align。主观与客观评估表明,本方法在音频质量、语义一致性及时间对齐方面均优于现有视频转音频模型。消融实验证明各模块有效性。音频样本见 https://y-ren16.github.io/STAV2A。

原文摘要 · Abstract (English)

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader applications. In this paper, we propose Semantic and Temporal Aligned Video-to-Audio (STA-V2A), an approach that enhances audio generation from videos by extracting both local temporal and global semantic video features and combining these refined video features with text as cross-modal guidance. To address the issue of information redundancy in videos, we propose an onset prediction pretext task for local temporal feature extraction and an attentive pooling module for global semantic feature extraction. To supplement the insufficient semantic information in videos, we propose a Latent Diffusion Model with Text-to-Audio priors initialization and cross-modal guidance. We also introduce Audio-Audio Align, a new metric to assess audio-temporal alignment. Subjective and objective metrics demonstrate that our method surpasses existing Video-to-Audio models in generating audio with better quality, semantic consistency, and temporal alignment. The ablation experiment validated the effectiveness of each module. Audio samples are available at https://y-ren16.github.io/STAV2A.

视频生成音频生成跨模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。