arXiv:2606.14141cs.SDcs.AI2026-06

让声音模型同时理解内容、位置和运动轨迹,提升音频问答能力。

Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

论文配图:Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources
图 1 · 摘自论文原文
  • 用时序化全景音编码器捕捉声音的语义与动态轨迹。
  • 在ST-AudioQA数据集上推理准确率显著优于静态与定位基线。
  • 适合做声音定位、多模态推理和智能听觉系统的研究者。

声音事件具有语义身份、位置和运动轨迹,但现有音频-语言模型通常将音频片段视为全局事件内容。相反,声音定位模型能追踪声源方向随时间变化,但语义覆盖有限,难以支持语言推理。为填补这一空白,我们构建了首个时空音频问答数据集ST-AudioQA,基于一阶球面谐波(FOA)渲染的静止与移动声源场景,每场景包含声源身份、活动状态、方向、距离及运动元数据,支持密集轨迹标注,并可提出关于声音内容、位置、运动方式及声源间关系的问题。我们进一步提出ST-Audio Encoder,一种时间分辨的FOA音频编码器,联合学习事件语义与源轨迹;以及ST-AudioLM,将编码器输出的音频标记连接至大语言模型实现时空音频问答。实验表明,该表示在语义-定位权衡上表现更优,推理性能显著超越静态空间与定位导向基线。

原文摘要 · Abstract (English)

Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, sound event localization models track source directions over time but offer limited semantic coverage for language reasoning. To address this gap, we introduce ST-AudioQA, a spatio-temporal audio QA dataset and benchmark built from first-order ambisonic (FOA) renderings of static and moving sound sources. Each scene provides source identity, activity, direction, distance, and motion metadata, enabling dense trajectory supervision and questions about what is sounding, where it is, how it moves, and how sources relate. We further propose ST-Audio Encoder, a time-resolved FOA audio encoder that learns event semantics together with source trajectories, and ST-AudioLM, which connects the audio tokens from the encoder to an LLM for spatio-temporal audio QA. Experiments show that this representation improves the semantic-localization tradeoff and yields stronger reasoning performance than static spatial and localization-oriented baselines.

音频理解时空建模多模态语音定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。