统一生成视频同步的语音、歌曲和音频,精度与效率俱佳。
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
- 用多模态扩散变换器联合训练视频-文本-音频数据,实现跨模态生成
- 生成8秒音频仅需1.91秒,唇形同步准确率达顶尖水平
- 适合影视配音、虚拟偶像、智能音视频合成等场景
我们提出AudioGen-Omni——一种基于多模态扩散变换器(MMDit)的统一方法,可生成与输入视频高度同步的高质量音频、语音及歌曲。该模型通过新型联合训练范式,无缝整合大规模视频-文本-音频语料库,具备在多模态输入下生成语义丰富、声学多样的音频能力,并适配多种音频生成任务。AudioGen-Omni采用统一的歌词-转录编码器,将歌曲与口语中的字符与音素编码为密集帧级表示。通过基于AdaLN的联合注意力机制结合相位对齐各向异性位置注入(PAAPI),仅对时序结构化模态施加旋转位置编码(RoPE),确保跨模态精准对齐。通过解冻所有模态并掩码缺失输入,克服了文本冻结范式的语义限制,实现有效跨模态条件生成。该方法显著提升音频质量、语义一致性与唇形同步精度,在文本到音频/语音/歌曲任务中达到当前最优性能。推理速度为每8秒音频仅需1.91秒,大幅优化效率与泛化能力。
原文摘要 · Abstract (English)
We present AudioGen-Omni - a unified approach based on multimodal diffusion transformers (MMDit), capable of generating high-fidelity audio, speech, and song coherently synchronized with the input video. AudioGen-Omni introduces a novel joint training paradigm that seamlessly integrates large-scale video-text-audio corpora, enabling a model capable of generating semantically rich, acoustically diverse audio conditioned on multimodal inputs and adaptable to a wide range of audio generation tasks. AudioGen-Omni employs a unified lyrics-transcription encoder that encodes graphemes and phonemes from both song and spoken inputs into dense frame-level representations. Dense frame-level representations are fused using an AdaLN-based joint attention mechanism enhanced with phase-aligned anisotropic positional infusion (PAAPI), wherein RoPE is selectively applied to temporally structured modalities to ensure precise and robust cross-modal alignment. By unfreezing all modalities and masking missing inputs, AudioGen-Omni mitigates the semantic constraints of text-frozen paradigms, enabling effective cross-modal conditioning. This joint training approach enhances audio quality, semantic alignment, and lip-sync accuracy, while also achieving state-of-the-art results on Text-to-Audio/Speech/Song tasks. With an inference time of 1.91 seconds for 8 seconds of audio, it offers substantial improvements in both efficiency and generality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。