arXiv:2509.09526eess.AScs.SD2025-09

按空间区域精准识别音频中的声音事件,提升定位感知能力。

Region-Specific Audio Tagging for Spatial Sound

  • 基于麦克风阵列的声场数据,区分不同角度或距离的声音来源。
  • 在仿真与真实数据集上验证方法有效,方向特征提升全向标签精度。
  • 适合需要空间感知的音频分析场景,如智能音箱、虚拟现实。

音频标注旨在为音频记录中的声音事件打标签。本文提出一种新任务——区域特定音频标注,针对由麦克风阵列录制的空间音频,对指定区域内的声音事件进行标注,该区域可定义为角度范围或距麦克风的距离。我们首先研究了频谱、空间和位置特征的不同组合性能;随后将先进的音频标注系统(如预训练音频神经网络PANNs和音频频谱变换器AST)扩展至该任务。在模拟与真实数据集上的实验表明该任务的可行性及所提方法的有效性。进一步实验显示,引入方向特征有助于提升全向音频标注性能。

原文摘要 · Abstract (English)

Audio tagging aims to label sound events appearing in an audio recording. In this paper, we propose region-specific audio tagging, a new task which labels sound events in a given region for spatial audio recorded by a microphone array. The region can be specified as an angular space or a distance from the microphone. We first study the performance of different combinations of spectral, spatial, and position features. Then we extend state-of-the-art audio tagging systems such as pre-trained audio neural networks (PANNs) and audio spectrogram transformer (AST) to the proposed region-specific audio tagging task. Experimental results on both the simulated and the real datasets show the feasibility of the proposed task and the effectiveness of the proposed method. Further experiments show that incorporating the directional features is beneficial for omnidirectional tagging.

音频标注空间音频麦克风阵列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。