让视频人物说话时声音和长相同步匹配参考源。
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
- 用独立的图像与音频控制模块,实现零样本联合生成。
- 对比学习提升音色与身份一致性,生成更自然。
- 适合需要真人级音画同步的数字人应用。
现有主流视频定制方法主要基于参考图像和文本提示生成身份一致的视频。得益于联合音视频生成技术的快速发展,本文提出更具挑战性的新任务:音视频同步定制,旨在同步定制视频身份与音频音色。具体而言,给定参考图像 $I^{r}$ 和参考音频 $A^{r}$,该任务要求生成保持参考图像身份、模仿参考音频音色,并可通过用户提供的文本提示自由指定口语内容的视频。为此,我们提出 OmniCustom,一个基于 DiT 的强大音视频定制框架,可零样本地同时遵循参考图像身份、音频音色和文本提示生成视频。本框架包含三项关键贡献:第一,通过嵌入在基础音视频生成模型自注意力层中的独立身份与音频 LoRA 模块,实现身份与音色控制;第二,引入对比学习目标,以条件预测流为正例、无参考条件流为负例,增强模型对身份与音色的保持能力;第三,我们在自建的大规模高质量音视频人像数据集上训练 OmniCustom。大量实验表明,OmniCustom 在生成音视频内容的身份与音色一致性方面优于现有方法。
原文摘要 · Abstract (English)
Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, this paper proposes a more compelling new task: sync audio-video customization, which aims to synchronously customize both video identity and audio timbre. Specifically, given a reference image $I^{r}$ and a reference audio $A^{r}$, this novel task requires generating videos that maintain the identity of the reference image while imitating the timbre of the reference audio, with spoken content freely specifiable through user-provided textual prompts. To this end, we propose OmniCustom, a powerful DiT-based audio-video customization framework that can synthesize a video following reference image identity, audio timbre, and text prompts all at once in a zero-shot manner. Our framework is built on three key contributions. First, identity and audio timbre control are achieved through separate reference identity and audio LoRA modules that operate through self-attention layers within the base audio-video generation model. Second, we introduce a contrastive learning objective alongside the standard flow matching objective. It uses predicted flows conditioned on reference inputs as positive examples and those without reference conditions as negative examples, thereby enhancing the model ability to preserve identity and timbre. Third, we train OmniCustom on our constructed large-scale, high-quality audio-visual human dataset. Extensive experiments demonstrate that OmniCustom outperforms existing methods in generating audio-video content with consistent identity and timbre fidelity. Project page: https://omnicustom-project.github.io/page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。