让音视频生成更同步,解决多模态对齐难题。
Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy
- 用跨任务协同训练缓解音视频噪声潜变量漂移问题。
- 通过全局-局部解耦模块提升时间风格对齐精度。
- 新设计的SyncCFG增强推理时的跨模态同步信号。
生成同步音视频内容是生成式AI的关键挑战,开源模型常面临音视频对齐不稳的问题。分析表明,这源于联合扩散过程中的三大根本问题:(1) 对应漂移,即并发演化的噪声潜变量阻碍稳定对齐学习;(2) 低效的全局注意力机制,无法捕捉细粒度时间线索;(3) 传统无分类器引导(CFG)的模态内偏倚,虽增强条件性却无助于跨模态同步。为此,我们提出Harmony框架,从机制上强制音视频同步。首先引入跨任务协同训练范式,利用音频驱动视频与视频驱动音频任务的强监督信号缓解漂移;其次设计全局-局部解耦交互模块,实现高效精准的时间-风格对齐;最后提出新型同步增强型无分类器引导(SyncCFG),在推理阶段显式分离并放大对齐信号。大量实验表明,Harmony达到新最优性能,显著优于现有方法,在生成保真度和细粒度音视频同步方面均有突破。
原文摘要 · Abstract (English)
The synthesis of synchronized audio-visual content is a key challenge in generative AI, with open-source models facing challenges in robust audio-video alignment. Our analysis reveals that this issue is rooted in three fundamental challenges of the joint diffusion process: (1) Correspondence Drift, where concurrently evolving noisy latents impede stable learning of alignment; (2) inefficient global attention mechanisms that fail to capture fine-grained temporal cues; and (3) the intra-modal bias of conventional Classifier-Free Guidance (CFG), which enhances conditionality but not cross-modal synchronization. To overcome these challenges, we introduce Harmony, a novel framework that mechanistically enforces audio-visual synchronization. We first propose a Cross-Task Synergy training paradigm to mitigate drift by leveraging strong supervisory signals from audio-driven video and video-driven audio generation tasks. Then, we design a Global-Local Decoupled Interaction Module for efficient and precise temporal-style alignment. Finally, we present a novel Synchronization-Enhanced CFG (SyncCFG) that explicitly isolates and amplifies the alignment signal during inference. Extensive experiments demonstrate that Harmony establishes a new state-of-the-art, significantly outperforming existing methods in both generation fidelity and, critically, in achieving fine-grained audio-visual synchronization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。