arXiv:2509.14666cs.SDcs.AI2025-09

让机器理解移动声源的空间位置和动态变化,实现听觉场景推理。

Spatial Audio Motion Understanding and Reasoning

  • 构建空间音频编码器,帧级检测多声源并估计方向与距离。
  • 引入语义对齐模型,使音频特征与文本嵌入匹配,提升泛化能力。
  • 结合大语言模型,支持对动态声场的复杂问题问答,适合语音智能系统研究者。

空间音频推理使机器能够通过理解事件及其空间属性来解析听觉场景。本文聚焦于移动声源的空间音频理解与推理。首先,提出一种空间音频编码器,可在帧级别检测多个重叠事件,并估计其方向(DoA)和源距离。为提升对未见事件的泛化能力,引入音频基础模型,通过交叉注意力机制将音频特征与语义类别文本嵌入对齐。其次,为回答涉及移动声源的复杂动态音频场景问题,将大语言模型(LLM)基于本模型提取的结构化空间属性进行条件化。最后,构建了一个空间音频运动理解与推理基准数据集,并验证了框架在该数据集上的性能,优于基线模型。

原文摘要 · Abstract (English)

Spatial audio reasoning enables machines to interpret auditory scenes by understanding events and their spatial attributes. In this work, we focus on spatial audio understanding with an emphasis on reasoning about moving sources. First, we introduce a spatial audio encoder that processes spatial audio to detect multiple overlapping events and estimate their spatial attributes, Direction of Arrival (DoA) and source distance, at the frame level. To generalize to unseen events, we incorporate an audio grounding model that aligns audio features with semantic audio class text embeddings via a cross-attention mechanism. Second, to answer complex queries about dynamic audio scenes involving moving sources, we condition a large language model (LLM) on structured spatial attributes extracted by our model. Finally, we introduce a spatial audio motion understanding and reasoning benchmark dataset and demonstrate our framework's performance against the baseline model.

音频理解空间推理大语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。