针对1秒内3D声音定位,提出滤波器与SCConv改进方案。
Enhancing 1-Second 3D SELD Performance with Filter Bank Analysis and SCConv Integration in CST-Former
- 用伽马音滤波器提取音频特征,提升短时精度
- 替换CST-Former中的卷积模块,F分数显著提升
- 适合低延迟实时语音定位场景应用
近期的语音事件检测与定位(SELD)研究多集中于5至10秒的长段落场景,虽提升基准性能,但缺乏真实应用所需的时序精细度。本文首次研究1秒短时窗口下的三维声音定位(3D SELD),建立面向实际应用的新基准。比较了Bark、Mel和伽马音(Gammatone)滤波器在音频特征提取中的表现,实验表明伽马音滤波器在该场景下整体准确率最高。进一步将领先架构CST-Former中的卷积模块替换为SCConv模块,在短时场景中实现可测量的F-score提升,验证了其在空间与通道特征建模上的优势。结果表明,该方法显著推进了3D SELD系统在低延迟条件下的实用化进程。
原文摘要 · Abstract (English)
Recent SELD research has predominantly focused on long-time segment scenarios (typically 5 to 10 seconds, occasionally 2 seconds), improving benchmark performance but lacking the temporal granularity needed for real-world applications. To bridge this gap, this paper investigates SELD with distance estimation (3D SELD) systems under short-time segments, specifically targeting a 1-second window, establishing a new baseline for practical 3D SELD applicability. We further explore the impact of different filter banks -- Bark, Mel, and Gammatone for audio feature extraction, and experimental results demonstrate that the Gammatone filter achieves the highest overall accuracy in this context. Finally, we propose replacing the convolutional modules within the CST-Former, a competitive SELD architecture, with the SCConv module. This adjustment yields measurable F-score gains in short-segment scenarios, underscoring SCConv's potential to improve spatial and channel feature representation. The experimental results highlight our approach as a significant step towards the real-world deployment of 3D SELD systems under low-latency constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。