用大视觉语言模型实现高保真视频配音,支持多模态协同生成。
DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
- 基于双视觉编码器与自回归架构,联合建模视频、音频与文本关系。
- 在多个基准上达到领先性能,生成音频与视频高度同步且质量优异。
- 开源关键数据集描述,助力后续研究高效评估与对比。
视频生成技术在视觉保真度上取得显著进展,但缺乏同步音频严重削弱沉浸感并限制实际应用。为解决该问题,本文提出一种基于大视觉语言模型(VLMs)的自回归音频生成框架DreamFoley,能够联合建模视频、音频与文本的序列交互关系。模型采用双视觉编码器,分别提取与音频对齐和文本对齐的视觉特征;引入带有延迟模式生成策略的残差向量量化音频分词器,在训练效率与音频质量间取得平衡;同时将无分类器引导机制引入VLM,提升生成音频质量。此外,构建高效的数据生产流水线,实现大规模音视频文本三元组的收集。大量实验验证了模型的有效性,在多个主流基准上表现优异。研究还公开了此前缺失的公共基准音频-视觉-文本描述,旨在为后续研究提供更便捷、高效的评估基础。
原文摘要 · Abstract (English)
Recent advances in video generation have achieved remarkable improvements in visual content fidelity. However, the absence of synchronized audio severely undermines immersive experience and restricts practical applications of these technologies. To address this challenge, several pioneering works have explored diffusion transformer architectures for generating plausible video-synchronized audio, including Kling-foley, HunyuanVideo-foley and Thinksound. Distinct from existing works, we introduce an autoregressive audio generation architecture (DreamFoley) that harnesses the capabilities of large vision-language models (VLMs) to jointly model sequential interactions among video, audio, and text modalities. Our approach features a dual-visual encoder module that effectively captures both audio-aligned and text-aligned visual features. Additionally, we employ a Residual Vector Quantization audio tokenizer with a delay-pattern generation scheme to balance the trade-off between training efficiency and audio quality. Moreover, we introduce the classifier-free guidance strategy into VLMs to bootstrap generated audio quality. Furthermore, we establish an efficient data production pipeline to scale audio-video-text triple collection. Finally, extensive experiments are conducted to validate the effectiveness of our model, achieving promising performance across popular benchmarks. We hope that the findings in this study provide a strong foundation for future video-to-audio generation research. We also release the previously missing audio-visual textual descriptions from the public benchmark, aiming to facilitate subsequent researchers in conducting more convenient and effective evaluations and comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。