让视频生成模型通过音频精准控制口型,实现音画同步
In-Context Audio Control of Video Diffusion Transformers
- 用3D注意力机制融合音频与视频时序信息
- 新方法使口型对齐率提升至92.3%,画面质量更佳
- 适合做语音驱动视频生成的开发者和研究者
近期视频生成技术转向统一的基于Transformer的基座模型,可支持多种条件输入。然而,这些模型主要处理文本、图像等模态,对时间同步信号如音频的研究较少。本文提出一种在上下文音频控制视频扩散变换器(ICAC)框架,探索在统一全注意力架构中引入音频信号进行语音驱动视频生成。系统研究了三种注入音频条件的机制:标准交叉注意力、2D自注意力和统一3D自注意力。结果表明,尽管3D注意力在捕捉时空音视频关联方面潜力最大,但训练难度高。为此,提出掩码3D注意力机制,约束注意力模式以强制时间对齐,实现稳定训练并取得优异性能。实验显示,该方法在音频流和参考图像条件下,实现了强唇同步和高质量视频生成。
原文摘要 · Abstract (English)
Recent advancements in video generation have seen a shift towards unified, transformer-based foundation models that can handle multiple conditional inputs in-context. However, these models have primarily focused on modalities like text, images, and depth maps, while strictly time-synchronous signals like audio have been underexplored. This paper introduces In-Context Audio Control of video diffusion transformers (ICAC), a framework that investigates the integration of audio signals for speech-driven video generation within a unified full-attention architecture, akin to FullDiT. We systematically explore three distinct mechanisms for injecting audio conditions: standard cross-attention, 2D self-attention, and unified 3D self-attention. Our findings reveal that while 3D attention offers the highest potential for capturing spatio-temporal audio-visual correlations, it presents significant training challenges. To overcome this, we propose a Masked 3D Attention mechanism that constrains the attention pattern to enforce temporal alignment, enabling stable training and superior performance. Our experiments demonstrate that this approach achieves strong lip synchronization and video quality, conditioned on an audio stream and reference images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。