提出一种高效生成长序列音视频的方法,解决多视角扩散模型的频谱失真问题。
Latent Swap Joint Diffusion for 2D Long-Form Latent Generation
- 采用帧级双向交换机制增强高频成分,避免频谱混淆
- 通过参考引导实现跨视图一致性,支持更长序列生成
- 比现有方法快2~20倍,适用于音频与全景图生成
本文提出一种名为SaFa的通用且高效的多视角联合扩散生成方法,用于生成无缝连贯的长序列音频与全景图。针对现有联合扩散方法在基于频谱的音频生成中出现的频谱混叠问题,通过对比梅尔频谱与RGB图像的VAE隐空间表示,发现其根源在于高频分量在频谱去噪过程中被平均算子过度抑制。为此,提出自环隐空间交换(Self-Loop Latent Swap),在相邻视图重叠区域实施帧级双向交换,利用子视图间的步进差异化轨迹自适应增强高频成分,防止频谱失真。为进一步提升非重叠区域的全局跨视图一致性,引入参考引导隐空间交换(Reference-Guided Latent Swap),通过中心化参考轨迹同步子视图扩散过程。通过优化交换时机与间隔,可在仅前向传播下实现跨视图相似性与多样性平衡。定量与定性实验表明,SaFa在使用U-Net与DiT模型的音频生成任务中显著优于现有联合扩散方法,甚至超越部分训练型方法,并具备良好的长序列扩展能力;在全景图生成中性能相当,速度提升2~20倍,模型泛化性更强。更多生成演示见https://swapforward.github.io/
原文摘要 · Abstract (English)
This paper introduces Swap Forward (SaFa), a modality-agnostic and efficient method to generate seamless and coherence long spectrum and panorama through latent swap joint diffusion across multi-views. We first investigate the spectrum aliasing problem in spectrum-based audio generation caused by existing joint diffusion methods. Through a comparative analysis of the VAE latent representation of Mel-spectra and RGB images, we identify that the failure arises from excessive suppression of high-frequency components during the spectrum denoising process due to the averaging operator. To address this issue, we propose Self-Loop Latent Swap, a frame-level bidirectional swap applied to the overlapping region of adjacent views. Leveraging stepwise differentiated trajectories of adjacent subviews, this swap operator adaptively enhances high-frequency components and avoid spectrum distortion. Furthermore, to improve global cross-view consistency in non-overlapping regions, we introduce Reference-Guided Latent Swap, a unidirectional latent swap operator that provides a centralized reference trajectory to synchronize subview diffusions. By refining swap timing and intervals, we can achieve a cross-view similarity-diversity balance in a forward-only manner. Quantitative and qualitative experiments demonstrate that SaFa significantly outperforms existing joint diffusion methods and even training-based methods in audio generation using both U-Net and DiT models, along with effective longer length adaptation. It also adapts well to panorama generation, achieving comparable performance with 2 $\sim$ 20 $\times$ faster speed and greater model generalizability. More generation demos are available at https://swapforward.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。