将通用音频分类模型升级为具备空间定位能力的声事件检测系统
From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection

- 用预训练音频标签模型搭配全向音场处理,实现语义到空间的精准迁移
- 通过分阶段神经架构搜索,找到最优的空间特征融合策略
- 在多个数据集上验证了跨域迁移能力,适合部署于资源受限场景
本报告研究如何将预训练的通用音频标签(GP-AT)模型扩展至具备空间定位能力的声事件定位与检测(SELD)。提出的AT2SELD框架结合预训练的AT骨干网络与紧凑的一阶全向音场(FOA)空间处理,采用逐轨声事件检测与笛卡尔方向角(DOA)估计、排列感知监督及校准机制。该框架揭示了语义音频先验在数据、计算和部署约束下支持定位感知场景分析的能力。通过有指导的多阶段神经架构搜索(NAS)构建:第一阶段表明基于幅度、相位和强度矢量(IVs)的谱域FOA描述符是语义到空间迁移最可靠的接口;第二阶段识别出早期残差空间编码为主要容量敏感组件,而后期轨道级抽象与递归平滑主要起精炼作用;第三阶段显示晚期交叉缝合能提升语义-空间交互,而早期融合成本更高且效果更差。诊断评估分析所选架构在类平衡、焦点损失、活动条件化DOA监督、阈值校准以及跨STARSS23、TAU2019、TAU-NIGENS2020和TAU-NIGENS2021数据集的表现。焦点损失提升活跃点表现,仅对活跃目标进行DOA监督可缓解非活跃目标主导问题,验证选择的阈值可在不替换空间学习的前提下恢复校准。跨数据集与理想活动分析表明,在TAU2019上具有强固定源定位能力,TAU NIGENS2021提供可迁移表征,而在STARSS23上表现有意义但不确定。总体而言,当嵌入空间感知架构并经过集成校准与部署导向优化时,GP-AT先验在SELD设计中展现出巨大潜力。
原文摘要 · Abstract (English)
This report investigates the extension of pretrained General-Purpose Audio Tagging (GP-AT) models toward spatially grounded Sound Event Localization and Detection (SELD). The proposed AT2SELD framework couples a pretrained AT backbone with compact First-Order Ambisonics (FOA) spatial processing, track-wise SED and Cartesian DOA estimation, permutation aware supervision, and calibration. It characterizes how semantic audio priors support localization-aware scene analysis under data, computation, and deployment constraints. The framework is developed through informed multi-stage Neural Architecture Search (NAS). Stage 1 shows that spectral FOA descriptors, based on magnitude, phase, and Intensity Vectors (IVs), provide the most reliable interface for semantic-to-spatial transfer. Stage 2 identifies early residual spatial encoding as the main capacity-sensitive component, while late track-wise abstraction and recurrent smoothing act mainly as refinement stages. Stage 3 shows that late cross-stitch coupling improves semantic-spatial interaction, whereas early fusion is costlier and less effective. Diagnostic evaluation analyzes the selected architecture under class balancing, focal loss, activity-conditioned DOA supervision, threshold calibration, and transfer across STARSS23, TAU2019, TAU-NIGENS2020, and TAU-NIGENS2021. Focal loss improves the activity point, active-only DOA supervision mitigates inactive target dominance, and validation-selected thresholds recover calibration without replacing spatial learning. Cross-dataset and oracle-activity analyses indicate strong fixed source localization on TAU2019, transferable representations from TAU NIGENS2021, and meaningful but uncertain behavior on STARSS23. Overall, GP-AT priors appear promising for SELD design when embedded in spatial-aware architectures and optimized through integrated calibration and deployment oriented strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。