arXiv:2412.18157cs.SDcs.AI2024-12被引 3

用文本语义引导生成更连贯、对齐更准的视频配乐。

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance

  • 通过文本标签持续提供语义引导,提升音视频时空对齐精度。
  • 引入帧级与时间适配器,融合高分辨率视觉特征与跨模态相似性。
  • 适合需要物理合理、连续音频的影视音效生成场景。

视频到音频(V2A)生成任务因在制作拟音音效方面的实用性受到关注。语义与时间条件被输入生成模型以指示声音事件及其发生时间。现有方法在存在动态画面的视频上面临挑战:时间条件不够精确,低分辨率语义条件加剧问题。为此,我们提出Smooth-Foley,一种利用全程文本标签进行语义引导的V2A生成模型,以增强音频的语义与时间对齐。训练两个适配器以利用预训练的文本到音频生成模型:帧适配器融合逐帧视频特征,时间适配器基于视觉帧与文本标签间的相似性整合时间条件。文本语义引导使生成音频实现更精确的音视频对齐。我们进行了广泛的定量与定性实验。结果表明,Smooth-Foley在连续声音与通用场景下均优于现有模型。在语义引导下,生成音频质量更高,且更符合物理规律。

原文摘要 · Abstract (English)

The video-to-audio (V2A) generation task has drawn attention in the field of multimedia due to the practicality in producing Foley sound. Semantic and temporal conditions are fed to the generation model to indicate sound events and temporal occurrence. Recent studies on synthesizing immersive and synchronized audio are faced with challenges on videos with moving visual presence. The temporal condition is not accurate enough while low-resolution semantic condition exacerbates the problem. To tackle these challenges, we propose Smooth-Foley, a V2A generative model taking semantic guidance from the textual label across the generation to enhance both semantic and temporal alignment in audio. Two adapters are trained to leverage pre-trained text-to-audio generation models. A frame adapter integrates high-resolution frame-wise video features while a temporal adapter integrates temporal conditions obtained from similarities of visual frames and textual labels. The incorporation of semantic guidance from textual labels achieves precise audio-video alignment. We conduct extensive quantitative and qualitative experiments. Results show that Smooth-Foley performs better than existing models on both continuous sound scenarios and general scenarios. With semantic guidance, the audio generated by Smooth-Foley exhibits higher quality and better adherence to physical laws.

音视频生成语义引导拟音时序对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。