arXiv:2506.00927cs.SDcs.AI2025-06ACL被引 5

用文字提示精准控制音频空间位置,实现沉浸式音效生成

In-the-wild Audio Spatialization with Flexible Text-guided Localization

  • 通过文本描述控制音频在三维空间中的定位,支持交互式调整
  • 在37.6万条模拟立体声数据上训练,实测生成效果优于现有方法
  • 结合大模型评估空间语义一致性,适合元宇宙与智能交互场景

为提升沉浸式体验,双耳音频在AR、VR和具身AI应用中提供了声音对象的空间感知。现有音频空间化方法虽可将单声道音频转换为双耳信号,但在复杂多对象用户交互环境中缺乏灵活可控性。为此,我们提出文本引导音频空间化(TAS)框架,利用灵活的文本提示,并从生成与理解双重角度评估模型。由于高质量大规模立体声数据稀缺,我们构建了包含376,000条模拟双耳音频样本的SpatialTAS数据集,以支持模型训练。模型学习基于三维空间位置和相对位置提示的双耳差异特征,辅以通道翻转音频增强。在模拟与真实录音数据集上均表现优异,验证了其卓越的泛化能力与精度。此外,我们基于Llama-3.1-8B开发评估模型,通过空间推理任务衡量生成双耳音频与文本提示之间的空间语义一致性。结果表明,文本提示可实现灵活交互控制,生成高质量且空间语义一致的双耳音频。数据集已公开于https://github.com/Alice01010101/TASU。

原文摘要 · Abstract (English)

To enhance immersive experiences, binaural audio offers spatial awareness of sounding objects in AR, VR, and embodied AI applications. While existing audio spatialization methods can generally map any available monaural audio to binaural audio signals, they often lack the flexible and interactive control needed in complex multi-object user-interactive environments. To address this, we propose a Text-guided Audio Spatialization (TAS) framework that utilizes flexible text prompts and evaluates our model from unified generation and comprehension perspectives. Due to the limited availability of premium and large-scale stereo data, we construct the SpatialTAS dataset, which encompasses 376,000 simulated binaural audio samples to facilitate the training of our model. Our model learns binaural differences guided by 3D spatial location and relative position prompts, augmented by flipped-channel audio. It outperforms existing methods on both simulated and real-recorded datasets, demonstrating superior generalization and accuracy. Besides, we develop an assessment model based on Llama-3.1-8B, which evaluates the spatial semantic coherence between our generated binaural audio and text prompts through a spatial reasoning task. Results demonstrate that text prompts provide flexible and interactive control to generate binaural audio with excellent quality and semantic consistency in spatial locations. Dataset is available at \href{https://github.com/Alice01010101/TASU}

音频空间化文本控制双耳音频沉浸式交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。