arXiv:2506.19774eess.AScs.AI2025-06被引 41

用多模态扩散模型生成与视频同步的高质量音效。

Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation

  • 采用多模态扩散变压器融合视频、音频与文本信息。
  • 在多个评估维度上达到公开模型最优表现,支持音效/语音/歌唱等场景。
  • 开源工业级评测基准Kling-Audio-Eval,适合音视频生成研究者使用。

我们提出Kling-Foley,一个大规模多模态视频到音频生成模型,可生成与视频内容同步的高质量音频。该模型引入多模态扩散变压器以建模视频、音频与文本之间的交互,并结合视觉语义表示模块和音视频同步模块,实现帧级别视频条件与潜在音频元素的对齐,提升语义一致性与音视频同步性。在文本条件协同下,可精准生成匹配视频的音效。此外,我们设计通用潜空间音频编码器,支持音效、语音、歌唱和音乐等多种场景的高质量建模。采用立体声渲染方法使合成音频具有空间感。为弥补开源基准数据集类型不全与标注不足的问题,我们还开源了工业级评测基准Kling-Audio-Eval。实验表明,基于流匹配目标训练的Kling-Foley在分布匹配、语义对齐、时间对齐和音频质量等方面均达到公开模型最新水平。

原文摘要 · Abstract (English)

We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce multimodal diffusion transformers to model the interactions between video, audio, and text modalities, and combine it with a visual semantic representation module and an audio-visual synchronization module to enhance alignment capabilities. Specifically, these modules align video conditions with latent audio elements at the frame level, thereby improving semantic alignment and audio-visual synchronization. Together with text conditions, this integrated approach enables precise generation of video-matching sound effects. In addition, we propose a universal latent audio codec that can achieve high-quality modeling in various scenarios such as sound effects, speech, singing, and music. We employ a stereo rendering method that imbues synthesized audio with a spatial presence. At the same time, in order to make up for the incomplete types and annotations of the open-source benchmark, we also open-source an industrial-level benchmark Kling-Audio-Eval. Our experiments show that Kling-Foley trained with the flow matching objective achieves new audio-visual SOTA performance among public models in terms of distribution matching, semantic alignment, temporal alignment and audio quality.

音视频生成扩散模型多模态音频合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。