用跨模态对齐提升视频生成音频的语义准确性。
FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
- 通过GRAM对齐视频、文本和音频嵌入,实现语义精准控制。
- 在Greatest Hits数据集上优于现有方法,音频与视频更匹配。
- 适合需要高语义一致性音效生成的研究者或开发者。
本文提出FoleyGRAM,一种基于跨模态对齐的视频到音频生成新方法。该方法利用格拉米安表示对齐度量(GRAM)对齐视频、文本与音频模态的嵌入表示,从而实现对音频生成过程的精确语义控制。核心采用扩散模型生成音频,并以GRAM对齐的嵌入与波形包络为条件,确保生成音频兼具语义丰富性与时序对齐性。在Greatest Hits数据集上的实验表明,使用GRAM对齐多模态编码器显著提升了生成音频与视频内容的语义一致性,推动了视频到音频合成的技术前沿。
原文摘要 · Abstract (English)
In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM leverages the Gramian Representation Alignment Measure (GRAM) to align embeddings across video, text, and audio modalities, enabling precise semantic control over the audio generation process. The core of FoleyGRAM is a diffusion-based audio synthesis model conditioned on GRAM-aligned embeddings and waveform envelopes, ensuring both semantic richness and temporal alignment with the corresponding input video. We evaluate FoleyGRAM on the Greatest Hits dataset, a standard benchmark for video-to-audio models. Our experiments demonstrate that aligning multimodal encoders using GRAM enhances the system's ability to semantically align generated audio with video content, advancing the state of the art in video-to-audio synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。