arXiv:2608.02474cs.CV2026-08中稿 · ACM MM 2026

用音频能量引导跨模态缓存,提升音频驱动视频生成效率

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation

论文配图:EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation
图 1 · 摘自论文原文
  • 以音频时频能量为关键锚点,动态更新潜在空间缓存
  • 在EMTD基准上实现2.46倍加速,保持音画一致性和生成质量
  • 适合追求高效音视频生成的开发者与研究者

音频驱动视频生成(A2V)在合成时间连贯且音画对齐的视频方面取得了显著进展,但其推理过程因扩散模型的迭代去噪而成本高昂。现有缓存方法主要利用视觉特征的时间冗余,忽略了A2V中音频驱动视觉生成带来的非均匀时间重要性。本文指出现有A2V缓存方法存在时序-语义与计算-存储不匹配两大问题。为此,提出EchoCache——一种基于能量引导的跨模态缓存框架。该框架利用音频时频能量作为显著性锚点,指导潜在空间缓存更新,并引入量化管理的动态时间步-潜在缓存机制,兼顾效率与内存优化。在主流A2V模型上的大量实验表明,EchoCache持续改善延迟-质量权衡,同时保持生成质量和音画一致性。特别地,在Wan2.2-S2V模型上,于EMTD基准下实现2.46倍加速,达到最优综合性能。代码已开源。

原文摘要 · Abstract (English)

Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its inference remains expensive due to the iterative denoising process of diffusion models. Existing caching methods mainly exploit temporal redundancy in visual features while overlooking the cross-modal alignment of A2V, where audio drives visual generation with highly non-uniform temporal importance. In this paper, we identify two levels of misalignment in existing A2V caching methods: temporal-semantic and computation-storage misalignment. To address them, we propose EchoCache, an energy-guided cross-modal caching framework for efficient A2V generation. EchoCache leverages audio time-frequency energy as a saliency anchor to guide latent-level cache updates and further introduces a dynamic timestep-latent caching mechanism with quantized cache management for joint efficiency and memory optimization. Extensive experiments on mainstream A2V models show that EchoCache consistently improves the latency-quality trade-off while preserving generation quality and audio-visual consistency. In particular, on Wan2.2-S2V over the EMTD benchmark, EchoCache achieves a 2.46x speedup with the best overall performance. Code is available at https://github.com/IF-LAB-PKU/EchoCache.

音视频生成缓存优化扩散模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。