arXiv:2606.03672cs.SDcs.MM2026-06被引 3

统一生成视频全音轨,语音音效音乐协同输出

Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation

论文配图:Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation
图 1 · 摘自论文原文
  • 共享潜在空间联合建模语音、音效与音乐生成
  • 混合音轨生成在听觉一致性与清晰度上显著提升
  • 适合影视后期、智能内容生成场景使用

当前统一音频生成模型虽能支持语音、音效和音乐等多任务,但多数仍局限于孤立的任务级合成。然而真实视频制作常需对同一视频同时生成包含语音、音效与音乐的完整音轨。我们提出Foley-Omni,一个统一的多模态音频生成模型,通过在共享潜在生成过程中联合建模语音、音效与音乐,实现从任务级合成到完整音轨生成的拓展。为支持训练与可复现评估,我们构建了视听数据整理流程,并引入V2ST-Bench基准,用于整体视频音轨生成评估。实验表明,Foley-Omni在单个任务上性能媲美专家系统,同时在混合音轨生成中提升了语音可懂度、音画一致性与感知质量。

原文摘要 · Abstract (English)

Recent unified audio generation models can support diverse tasks across speech, sound effects, and music, but most of them still focus on isolated task-level synthesis. However, real video production often requires multiple components of a complete audio track to be generated jointly and consistently for the same video. We present Foley-Omni, a unified multimodal audio generation model that extends isolated task-level synthesis to complete video soundtrack generation by jointly modeling speech, sound effects, and music within a shared latent generation process. To support training and reproducible evaluation, we develop an audiovisual data curation pipeline and introduce V2ST-Bench, a benchmark for holistic video soundtrack generation evaluation. Experiments show that Foley-Omni achieves competitive performance with expert systems on individual synthesis tasks, while improving speech intelligibility, audiovisual consistency and perceptual quality for mixed soundtrack generation.

音频生成多模态音轨合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。