统一视频生成音频,支持多模态控制与精准语义。
FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision

- 通过条件注入实现音视频协同生成与音效扩展。
- 在三个数据集上表现优于现有方法,实现帧级时间对齐。
- 适合需要精细音频控制的影视音效制作场景。
我们提出FoleyGenEx,一种统一的视频到音频(VTA)框架,集成多模态控制、帧级时间对齐和细粒度语义,实现多样化任务下的同步、多功能音频合成。现有VTA方法或具备多模态控制但时间对齐弱,或对齐强但缺乏参考音频条件和语义精度。FoleyGenEx通过三项核心创新填补这一空白:基于条件注入的音频控制式VTA与音效扩展机制,多模态动态掩码策略以保持训练同步性,以及结合信号处理与大语言模型的副词增强数据增广算法,提升文本监督的语义细腻度。在AudioCaps、VGGSound和Greatest Hits数据集上的实验表明,其可控VTA性能优于现有方法。演示样本可访问https://foleygenex.github.io/FoleyGenEx。
原文摘要 · Abstract (English)
We present FoleyGenEx, a unified video-to-audio (VTA) framework integrating multi-modal control, frame-level temporal alignment, and fine-grained semantics, enabling synchronized, versatile audio synthesis for diverse tasks. Existing VTA methods either have multi-modal control but weak temporal alignment or strong alignment but lack reference audio conditioning and semantic precision. FoleyGenEx fills this gap via three core innovations: a conditional injection mechanism for audio-controlled VTA and Foley extension, a multi-modal dynamic masking strategy preserving training synchronization, and an adverb-based data augmentation algorithm leveraging signal processing and large language models to enhance textual supervision with nuanced semantics. Experiments on AudioCaps, VGGSound, and Greatest Hits demonstrate its competitive controllable VTA performance against existing methods. Demo samples are available at https://foleygenex.github.io/FoleyGenEx.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。