让大模型听懂声音在哪,如何分布,还能判断场景是否合理。
The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models
- 用物理模拟的全景声数据训练模型,显式学习声音位置信息。
- 在多任务测试中达到70.8%整体准确率,场景级问答达79.76%。
- 适合研究音频理解、空间感知或需可解释推理的AI开发者。
大型音频-语言模型在识别音频内容上进展迅速,但缺乏对空间音频-语言理解的清晰任务接口。模型还需判断声音事件的位置、语义与空间属性的关联、多个对象的空间布局,以及场景级回答是否物理合理。本文提出音频场景分析(ASA)这一三层次问题:原子感知、关系整合与认知推理。为此设计了「世界非单声道」(TWNM)框架,通过物理基础的一阶全向声(FOA)模拟实现可控监督,从多通道音频中学习槽正则化空间表征,并融合语义音频特征,采用渐进式课程训练,最终以元数据衍生答案和辅助格式/证据奖励进行偏好优化。为实现ASA评估,构建基于场景元数据的受控基准,涵盖定位、属性绑定、空间比较、场景推断与反事实推理。TWNM在该基准上整体准确率达70.8%,空间族任务66.4%,混合层级三类场景问答79.76%。同时通过显式标签审计单声道与双声道参考系统,验证其在输入、训练接口与输出格式上的差异。结果表明,明确的ASA层级、FOA条件化空间表征与元数据驱动训练,能实现可控且可审计的空间音频-语言推理,STARSS23提供有限真实录音诊断支持。
原文摘要 · Abstract (English)
Large audio-language models have made rapid progress in recognizing what is present in an audio clip, but spatial audio-language understanding still lacks a clear task interface. A model must also decide where sound events occur, which semantic and spatial attributes belong to the same auditory object, how multiple objects are arranged, and whether a scene-level answer is physically plausible. We formalize this capability as audio scene analysis (ASA), a three-level problem spanning atomic perception, relational integration, and cognitive reasoning. We propose The World is Not Mono (TWNM), a framework that equips audio-language models with explicit spatial evidence. TWNM uses physically grounded First-Order Ambisonics (FOA) simulation for controllable supervision, learns slot-regularized spatial representations from multichannel audio, fuses them with semantic audio features, and trains with a progressive curriculum ending in preference optimization over metadata-derived answers and auxiliary format/evidence rewards. To operationalize ASA, we build a controlled benchmark from scene metadata, covering localization, attribute binding, spatial comparison, scene abduction, and counterfactual reasoning. On this benchmark, TWNM achieves 70.8% overall accuracy, 66.4% on spatial-family tasks, and 79.76% on mixed L3 scene-level multiple-choice QA. We also audit monaural and binaural reference systems as diagnostic references with explicit audit labels, since they differ in spatial input, training interface, and output format. The supported claim is that a clearly defined ASA hierarchy, FOA-conditioned spatial representations, and metadata-grounded training enable controlled, auditable spatial audio-language reasoning, with STARSS23 providing a limited real-recording diagnostic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。