提出两种混响特征,提升声音事件三维定位中距离估计精度。
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
- 用直达声与混响比(DRR)和自相关捕捉早期反射构造新特征
- 在STARSS23数据集上实现当前最优距离估计性能
- 适用于麦克风阵列和声学球面格式,适合三维声音定位研究者
声音事件定位与检测(SELD)需在时间维度上预测活跃声事件类别并估计其位置。传统定位多视为到达方向估计,忽略声源距离。近期研究将SELD扩展至三维空间,引入距离估计以实现三维定位(3D SELD)。然而现有方法缺乏专为距离估计设计的输入特征。本文提出两种基于混响的新特征:一种利用直达声与混响比(DRR),另一种通过信号自相关捕捉早期反射。我们在STARSS23数据集上对这些特征进行广泛评估,将其与经典SELD特征结合,用于声事件检测(SED)与到达方向估计(DOAE),并在不同网络架构下测试。所提特征适用于全向麦克风阵列(FOA)与普通麦克风阵列(MIC)格式,在距离估计方面达到当前最优表现,显著提升整体3D SELD性能。
原文摘要 · Abstract (English)
Sound event localization and detection (SELD) involves predicting active sound event classes over time while estimating their positions. The localization subtask in SELD is usually treated as a direction of arrival estimation problem, ignoring source distance. Only recently, SELD was extended to 3D by incorporating distance estimation, enabling the prediction of sound event positions in 3D space (3D SELD). However, existing methods lack input features specifically designed for distance estimation. We address this gap by introducing two novel reverberation-based feature formats: one using the direct-to-reverberant ratio (DRR) and another leveraging signal autocorrelation to capture early reflections. We extensively evaluate and benchmark these features on the STARSS23 dataset, combining them with established SELD features for sound event detection (SED) and direction-of-arrival estimation (DOAE), and testing across different network architectures. Our proposed features, applicable to both FOA and MIC formats, achieve state-of-the-art distance estimation, enhancing overall 3D SELD performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。