用文本控制立体声场景,实现逼真空间音频生成
Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation
- 构建仿真数据集BEWO-1M,支持多声源动态描述
- 提出SpatialSonic模型,实现空间方向精准引导
- 首个支持跨模态(文/图)的空间音频生成方法
扩散模型在单通道音频生成中表现优异,但立体声生成面临场景复杂、空间控制难的问题。本文首次尝试解决这一挑战,构建大规模仿真数据集BEWO-1M,包含丰富声景与动态多源描述,并通过检索获取配对的图像与立体声。针对现有模型生成空间音频随机模糊的问题,提出SpatialSonic模型,采用空间感知编码器与方位状态矩阵,提供精确空间引导。实验表明,该方法能生成符合物理规律、可控制且沉浸感强的空间音频,在模拟与真实数据上均优于现有方法。
原文摘要 · Abstract (English)
Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. Controlling stereo audio with spatial contexts remains challenging due to high data costs and unstable generative models. To the best of our knowledge, this work represents the first attempt to address these issues. We first construct a large-scale, simulation-based, and GPT-assisted dataset, BEWO-1M, with abundant soundscapes and descriptions even including moving and multiple sources. Beyond text modality, we have also acquired a set of images and rationally paired stereo audios through retrieval to advance multimodal generation. Existing audio generation models tend to generate rather random and indistinct spatial audio. To provide accurate guidance for Latent Diffusion Models, we introduce the SpatialSonic model utilizing spatial-aware encoders and azimuth state matrices to reveal reasonable spatial guidance. By leveraging spatial guidance, our model not only achieves the objective of generating immersive and controllable spatial audio from text but also extends to other modalities as the pioneer attempt. Finally, under fair settings, we conduct subjective and objective evaluations on simulated and real-world data to compare our approach with prevailing methods. The results demonstrate the effectiveness of our method, highlighting its capability to generate spatial audio that adheres to physical rules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。