arXiv:2412.20378cs.CVcs.MM2024-12中稿 · ed被引 11

用多模态条件与音量控制生成高保真视频配乐

Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control

  • 结合文本、音频和像素级视觉提示,实现精细音效合成
  • 支持60秒内任意时长的44.1kHz立体声输出,优于现有模型
  • 引入LUFS嵌入,可精确控制各声道音量变化,适配影视后期

视频到音频(V2A)生成利用仅含视觉信息的视频来生成与场景匹配的真实声音。然而,现有模型在生成音频的细粒度控制方面存在不足,尤其是在音量变化和多模态条件融合上。为此,我们提出Tri-Ergon,一个基于扩散模型的V2A方法,结合文本、听觉和像素级视觉提示,实现语义丰富且细节精准的音频合成。此外,我们引入了相对全尺度(LUFS)的音量嵌入,可对各音频通道随时间的音量变化进行精确手动控制,有效解决真实影视配音(Foley)流程中视频与音频间的复杂关联。Tri-Ergon能生成最高达60秒、采样率为44.1 kHz的高保真立体声音频片段,显著优于当前主流V2A方法通常仅生成固定时长单声道音频的局限。

原文摘要 · Abstract (English)

Video-to-audio (V2A) generation utilizes visual-only video features to produce realistic sounds that correspond to the scene. However, current V2A models often lack fine-grained control over the generated audio, especially in terms of loudness variation and the incorporation of multi-modal conditions. To overcome these limitations, we introduce Tri-Ergon, a diffusion-based V2A model that incorporates textual, auditory, and pixel-level visual prompts to enable detailed and semantically rich audio synthesis. Additionally, we introduce Loudness Units relative to Full Scale (LUFS) embedding, which allows for precise manual control of the loudness changes over time for individual audio channels, enabling our model to effectively address the intricate correlation of video and audio in real-world Foley workflows. Tri-Ergon is capable of creating 44.1 kHz high-fidelity stereo audio clips of varying lengths up to 60 seconds, which significantly outperforms existing state-of-the-art V2A methods that typically generate mono audio for a fixed duration.

视频生成音频合成扩散模型音量控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。