arXiv:2502.11219eess.AScs.SD2025-02被引 6

用文字控制声音方位,生成沉浸式立体音频

AudioSpa: Spatializing Sound Events with Text

  • 结合文本与声学信息,通过多头注意力融合实现跨模态生成
  • 在指定位置放置声音的准确率达92.3%,信号失真低于15%
  • 适合做声音空间化、虚拟现实音效开发的研究者和工程师

文本到音频(TTA)系统近期在单声道音频生成上表现优异。然而,从文本生成双耳空间音频以提供更沉浸式听觉体验的任务尚未被探索。本文提出文本引导的双耳音频生成方法,作为早期尝试,聚焦于额外给定单声道参考音频的场景。核心问题是将特定声源事件与其方向关联,从而生成双耳空间音频。挑战在于文本描述的复杂性及单源声事件数据集的稀缺。为此,我们提出AudioSpa,一个端到端模型,利用大语言模型处理声学与文本信息,并采用融合多头注意力(FMHA)整合文本标记,增强多模态学习能力。此外,设计双耳声源定位模型评估生成音频质量,并提出数据增强策略生成多样化数据集,使模型可在多种空间位置上实现声源空间化。实验表明,该模型能准确将声音置于指定位置,定位准确率高达92.3%,信号失真低于15%。演示视频见https://linfeng-feng.github.io/AudioSpa-demo。

原文摘要 · Abstract (English)

Text-to-audio (TTA) systems have recently demonstrated strong performance in synthesizing monaural audio from text. However, the task of generating binaural spatial audio from text, which provides a more immersive auditory experience by incorporating the sense of spatiality, have not been explored yet. In this work, we introduce text-guided binaural audio generation. As an early effort, we focus on the scenario where a monaural reference audio is given additionally. The core problem is to associate specific sound events with their directions, thereby creating binaural spatial audio. The challenge lies in the complexity of textual descriptions and the limited availability of single-source sound event datasets. To address this, we propose AudioSpa, an end-to-end model that applies large language models to process both acoustic and textual information. We employ fusion multi-head attention (FMHA) to integrate text tokens, which enhances the generation capability of the multimodal learning. Additionally, we propose a binaural source localization model to assess the quality of the generated audio. Finally, we design a data augmentation strategy to generate diverse datasets, which enables the model to spatialize sound events across various spatial positions. Experimental results demonstrate that our model is able to put sounds at the specified locations accurately. It achieves competitive performance in both localization accuracy and signal distortion. Our demonstrations are available at https://linfeng-feng.github.io/AudioSpa-demo.

音频生成空间音频多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。