arXiv:2411.17698cs.CVcs.MM2024-11CVPR被引 62

用视频+文本/音频控制,生成高质量且符合创意的音效

Video-Guided Foley Sound Generation with Multimodal Controls

  • 结合视频、文本和参考音效进行多模态条件控制
  • 支持48kHz全频带音效生成,可产生活泼或夸张的创意音效
  • 在真实与幻想音效间灵活切换,适合影视音效设计人员

为解决视频音效生成中需艺术化处理且控制灵活的问题,我们提出MultiFoley模型,实现基于视频的音效生成,并支持文本、音频及视频的多模态条件输入。给定无声视频和文本提示,该模型可生成干净音效(如无风噪的滑板轮声)或富有创意的音效(如狮子吼声变为猫叫)。同时支持从音效库或部分视频中选取参考音频作为条件。模型通过互联网视频数据(含低质音频)与专业音效录音联合训练,实现48kHz全频带高质量音频生成。自动化评估与人工评测均表明,MultiFoley能有效生成与视频同步的高质量音效,优于现有方法。

原文摘要 · Abstract (English)

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce MultiFoley, a model designed for video-guided sound generation that supports multimodal conditioning through text, audio, and video. Given a silent video and a text prompt, MultiFoley allows users to create clean sounds (e.g., skateboard wheels spinning without wind noise) or more whimsical sounds (e.g., making a lion's roar sound like a cat's meow). MultiFoley also allows users to choose reference audio from sound effects (SFX) libraries or partial videos for conditioning. A key novelty of our model lies in its joint training on both internet video datasets with low-quality audio and professional SFX recordings, enabling high-quality, full-bandwidth (48kHz) audio generation. Through automated evaluations and human studies, we demonstrate that MultiFoley successfully generates synchronized high-quality sounds across varied conditional inputs and outperforms existing methods. Please see our project page for video results: https://ificl.github.io/MultiFoley/

音效生成多模态视频控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。