arXiv:2604.04348cs.SDcs.CV2026-04被引 2

让视频+文字生成完整音景,包含可见/不可见声音与人声。

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

  • 基于流匹配的扩散模型,三路注意力同步处理可见声、不可见声和语音。
  • 在超过1000个样本上优于现有方法,人评与客观指标均领先。
  • 适合影视音效合成、虚拟场景生成等需要全维度音频的应用。

本文提出通用整体音频生成(UniHAGen)任务,旨在合成包含可见与不可见声音的完整听觉场景,覆盖环境事件、乐器演奏和人类语音等多领域。现有视频条件音频生成模型多仅关注可见声源,而近期联合文本-视频生成模型虽涵盖整体音景,但不支持语音生成。为此,我们提出OmniSonic,一种基于流匹配的扩散框架,联合视频与文本输入,采用TriAttn-DiT架构实现三重跨注意力机制,分别处理可见环境声、不可见环境声和语音条件,并通过混合专家(MoE)门控机制自适应调节三者贡献。此外,构建了包含千余样本的UniHAGen-Bench基准,涵盖三类典型语音-环境场景。大量实验表明,OmniSonic在客观指标与人工评估中均显著优于当前最优方法,为通用整体音频生成建立强基准。

原文摘要 · Abstract (English)

In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on-screen and off-screen sounds across diverse domains (e.g., ambient events, musical instruments, and human speech). Prior video-conditioned audio generation models typically focus on producing on-screen environmental sounds that correspond to visible sounding events, neglecting off-screen auditory events. While recent holistic joint text-video-to-audio generation models aim to produce auditory scenes with both on- and off-screen sound but they are limited to non-speech sounds, lacking the ability to generate or integrate human speech. To overcome these limitations, we introduce OmniSonic, a flow-matching-based diffusion framework jointly conditioned on video and text. It features a TriAttn-DiT architecture that performs three cross-attention operations to process on-screen environmental sound, off-screen environmental sound, and speech conditions simultaneously, with a Mixture-of-Experts (MoE) gating mechanism that adaptively balances their contributions during generation. Furthermore, we construct UniHAGen-Bench, a new benchmark with over one thousand samples covering three representative on/off-screen speech-environment scenarios. Extensive experiments show that OmniSonic consistently outperforms state-of-the-art approaches on both objective metrics and human evaluations, establishing a strong baseline for universal and holistic audio generation. Project page: https://weiguopian.github.io/OmniSonic_webpage/

音频生成多模态语音合成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。