arXiv:2508.16930eess.AScs.CV2025-08被引 46

让视频生成声音更真实,精准匹配画面动态和语义。

HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

  • 构建10万小时多模态数据集,自动标注提升训练质量。
  • 通过自监督音频特征对齐,显著提升音质与生成稳定性。
  • 融合视觉、文本与音频的扩散模型,适合影视音效生成场景。

近期视频生成技术已能产出高度逼真的视觉内容,但缺乏同步音频严重削弱沉浸感。为解决视频到音频生成中的多模态数据稀缺、模态不平衡及音频质量有限等关键问题,我们提出 HunyuanVideo-Foley,一个端到端的文本-视频-音频生成框架,可精确生成与视觉动态和语义上下文对齐的高保真音频。本方法包含三项核心创新:(1) 通过自动化标注构建可扩展的数据管道,整合100k小时多模态数据;(2) 采用自监督音频特征进行表征对齐,引导潜在扩散训练,有效提升音频质量与生成稳定性;(3) 设计新型多模态扩散变压器,通过双流音视频融合与联合注意力机制缓解模态竞争,并利用交叉注意力注入文本语义。全面评估表明,HunyuanVideo-Foley 在音频保真度、视觉语义对齐、时间对齐和分布匹配方面均达到当前最优水平。演示页面见:https://szczesnys.github.io/hunyuanvideo-foley/。

原文摘要 · Abstract (English)

Recent advances in video generation produce visually realistic content, yet the absence of synchronized audio severely compromises immersion. To address key challenges in video-to-audio generation, including multimodal data scarcity, modality imbalance and limited audio quality in existing methods, we propose HunyuanVideo-Foley, an end-to-end text-video-to-audio framework that synthesizes high-fidelity audio precisely aligned with visual dynamics and semantic context. Our approach incorporates three core innovations: (1) a scalable data pipeline curating 100k-hour multimodal datasets through automated annotation; (2) a representation alignment strategy using self-supervised audio features to guide latent diffusion training, efficiently improving audio quality and generation stability; (3) a novel multimodal diffusion transformer resolving modal competition, containing dual-stream audio-video fusion through joint attention, and textual semantic injection via cross-attention. Comprehensive evaluations demonstrate that HunyuanVideo-Foley achieves new state-of-the-art performance across audio fidelity, visual-semantic alignment, temporal alignment and distribution matching. The demo page is available at: https://szczesnys.github.io/hunyuanvideo-foley/.

音频生成扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。