通过关注音视频同步关键区域,提升音视频生成效率。
Efficient Audio-Visual Generation via Synchrony-Aware Cross-Modal Sparse Attention

- 基于音视频交叉注意力的结构化交互模式,识别关键区域。
- 保留同步关键区域的高精度计算,其余部分稀疏化处理。
- 在提速的同时保持音视频质量与同步性,适合高效生成场景。
近期的音视频生成模型可在统一扩散过程中合成同步的视频与声音,但推理成本仍高,因长视频标记序列需在去噪步骤中重复进行注意力计算。已有加速技术包括低比特量化、注意力稀疏化和特征缓存,但这些方法原本针对视频生成设计,直接应用于音视频模型时忽略了音视频分支间的交互,可能破坏音视频同步。本文提出一种面向音视频生成的同步感知加速框架。关键观察发现,双向音视频交叉注意力揭示了两个分支间的结构化交互,高响应值常集中在少数与声音相关的视觉和时间区域。基于此交互模式,我们引入保护性稀疏注意力策略,在保留同步关键标记的高保真计算的同时,对冗余注意力交互进行稀疏化。通过在加速过程中显式考虑跨模态依赖,本方法在提升推理效率的同时,维持了视频质量、音频质量及音视频同步性。
原文摘要 · Abstract (English)
Recent audio-visual generation models can synthesize synchronized video and sound in a unified diffusion process, but their inference cost remains high because long video token sequences require repeated attention computation across denoising steps. A variety of acceleration techniques have been developed for video generation models, including low-bit quantization, attention sparsification, and feature caching. However, since these methods are originally designed for video generation, directly applying them to audio-visual models overlooks the interactions between the audio and video branches and may therefore disrupt audio-video synchronization. We present a synchronization-aware acceleration framework for efficient audio-visual generation. Our key observation is that bidirectional audio-video cross-attention reveals structured interactions between the two branches, with high responses often concentrated on a few sound-related visual and temporal regions. Guided by this interaction pattern, we introduce a protected sparse attention strategy that preserves high-fidelity computation for synchronization-critical tokens while sparsifying redundant attention interactions. By explicitly accounting for cross-modal dependence during acceleration, our method improves inference efficiency while keeping video quality, audio quality, and audio-video synchronization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。