统一框架实现人像音视频的精准可控生成,支持多角色多声音解耦控制。
DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
- 采用对称条件扩散变换器,融合多源输入信号实现联合建模
- 在多人场景下实现身份与音色解耦,准确率显著优于现有方法
- 适合需要高可控性的人像音视频生成研究与工业应用
基础模型的进步推动了音视频联合生成的发展。然而,现有方法通常将参考驱动音视频生成(R2AV)、视频编辑(RV2AV)和音频驱动视频动画(RA2V)视为独立任务。同时,在单一框架内实现多角色身份与语音音色的精确、解耦控制仍是未解难题。本文提出DreamID-Omni,一个可控制的人像音视频生成统一框架。我们设计了对称条件扩散变换器,通过对称条件注入机制整合异构条件信号。为解决多人场景中普遍存在的身份-音色绑定失败和说话人混淆问题,提出双层级解耦策略:信号级同步旋转位置编码(Synchronized RoPE)确保注意力空间刚性绑定,语义级结构化描述(Structured Captions)建立显式属性-主体映射。此外,设计多任务渐进训练方案,利用弱约束生成先验正则化强约束任务,防止过拟合并协调不同目标。大量实验表明,DreamID-Omni在视频、音频及视听一致性方面均达到全面领先性能,甚至超越部分领先商业模型。代码将公开发布,以缩小学术研究与商用应用之间的差距。
原文摘要 · Abstract (English)
Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video editing (RV2AV) and audio-driven video animation (RA2V) as isolated objectives. Furthermore, achieving precise, disentangled control over multiple character identities and voice timbres within a single framework remains an open challenge. In this paper, we propose DreamID-Omni, a unified framework for controllable human-centric audio-video generation. Specifically, we design a Symmetric Conditional Diffusion Transformer that integrates heterogeneous conditioning signals via a symmetric conditional injection scheme. To resolve the pervasive identity-timbre binding failures and speaker confusion in multi-person scenarios, we introduce a Dual-Level Disentanglement strategy: Synchronized RoPE at the signal level to ensure rigid attention-space binding, and Structured Captions at the semantic level to establish explicit attribute-subject mappings. Furthermore, we devise a Multi-Task Progressive Training scheme that leverages weakly-constrained generative priors to regularize strongly-constrained tasks, preventing overfitting and harmonizing disparate objectives. Extensive experiments demonstrate that DreamID-Omni achieves comprehensive state-of-the-art performance across video, audio, and audio-visual consistency, even outperforming leading proprietary commercial models. We will release our code to bridge the gap between academic research and commercial-grade applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。