arXiv:2409.14709eess.AScs.SD2024-09被引 8

用帧级语义提升视频生成音频的时序对齐精度

Video-to-Audio Generation with Fine-grained Temporal Semantics

  • 基于潜在扩散模型,融合帧级语义信息增强音频时序对齐
  • 在客观与主观评估中均实现更高质量音频与精细对齐
  • 适合关注音视频同步生成的AI创作者与研究者

随着AIGC的发展,视频生成在学术界和工业界引发广泛关注(如Sora)。然而,生成与视频严格时序对齐的音频仍具挑战性,尤其因视频中包含复杂的语义信息。受文本到音频(TTA)生成成功启发,本文首次探索基于潜在扩散模型(LDM)的视频到音频(VTA)生成框架。初步实验表明LDM在VTA任务中具有巨大潜力,但仍存在时序对齐不足的问题。为此,本文提出利用近期流行的“接地分割任意模型”(Grounding SAM)提取视频帧中的细粒度语义信息,以增强VTA模型的时序对齐能力。大量实验结果表明,所提方法在客观与主观评价指标上均表现优异,实现了更优的音频质量和细粒度时序对齐。

原文摘要 · Abstract (English)

With recent advances of AIGC, video generation have gained a surge of research interest in both academia and industry (e.g., Sora). However, it remains a challenge to produce temporally aligned audio to synchronize the generated video, considering the complicated semantic information included in the latter. In this work, inspired by the recent success of text-to-audio (TTA) generation, we first investigate the video-to-audio (VTA) generation framework based on latent diffusion model (LDM). Similar to latest pioneering exploration in VTA, our preliminary results also show great potentials of LDM in VTA task, but it still suffers from sub-optimal temporal alignment. To this end, we propose to enhance the temporal alignment of VTA with frame-level semantic information. With the recently popular grounding segment anything model (Grounding SAM), we can extract the fine-grained semantics in video frames to enable VTA to produce better-aligned audio signal. Extensive experiments demonstrate the effectiveness of our system on both objective and subjective evaluation metrics, which shows both better audio quality and fine-grained temporal alignment.

音视频生成扩散模型时序对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。