arXiv:2511.06606eess.AScs.AI2025-11被引 4

让语音模型具备空间感知能力,能判断声音方向与远近。

SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models

  • 用FOA编码器提取声源方向、高度和距离等空间特征。
  • 在SPUR数据集上微调后,空间问答准确率显著提升。
  • 无需大改模型结构,适合想加入空间理解的开发者。

空间感知是听觉智能的核心,有助于精准理解真实声学场景,推动人类级环境感知。尽管近期大型音频-语言模型(LALMs)在复杂音频推理方面表现强劲,但多数仅处理单声道输入,缺乏对方向、仰角和距离等空间线索的捕捉能力。本文提出SPUR,一种轻量级、即插即用的方法,通过最小架构改动为LALMs注入空间感知能力。SPUR包含:(i) 一阶全向声场(FOA)编码器,将(W, X, Y, Z)四通道信号映射为旋转感知、以听者为中心的空间特征,并通过多模态适配器融入目标模型;(ii) SPUR-Set,一个结合开源FOA录音与受控模拟的时空问答数据集,强调相对方向、仰角、距离及重叠关系的监督推理。在SPUR-Set上微调模型后,空间问答与多说话人归属任务性能持续提升,同时保持通用音频理解能力。SPUR提供了一种简单有效的方案,可将单声道LALMs升级为具备空间意识的模型。大量消融实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Spatial perception is central to auditory intelligence, enabling accurate understanding of real-world acoustic scenes and advancing human-level perception of the world around us. While recent large audio-language models (LALMs) show strong reasoning over complex audios, most operate on monaural inputs and lack the ability to capture spatial cues such as direction, elevation, and distance. We introduce SPUR, a lightweight, plug-in approach that equips LALMs with spatial perception through minimal architectural changes. SPUR consists of: (i) a First-Order Ambisonics (FOA) encoder that maps (W, X, Y, Z) channels to rotation-aware, listener-centric spatial features, integrated into target LALMs via a multimodal adapter; and (ii) SPUR-Set, a spatial QA dataset combining open-source FOA recordings with controlled simulations, emphasizing relative direction, elevation, distance, and overlap for supervised spatial reasoning. Fine-tuning our model on the SPUR-Set consistently improves spatial QA and multi-speaker attribution while preserving general audio understanding. SPUR provides a simple recipe that transforms monaural LALMs into spatially aware models. Extensive ablations validate the effectiveness of our approach.

空间音频多模态语音理解听觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。