arXiv:2504.12870eess.AS2025-04被引 2

用多维注意力提升真实场景下声音事件定位与检测精度

CST-former: Multidimensional Attention-based Transformer for Sound Event Localization and Detection in Real Scenes

  • 设计多尺度局部嵌入模块,捕捉多时频尺度信息
  • 在STARSS22/23数据集上达到领先性能,无需外部数据
  • 适用于小样本训练的微调与后处理技术

声音事件定位与检测(SELD)旨在利用多通道音频信号对声音事件进行分类并识别其到达方向(DoA)。为有效实现分类与定位,本文提出通道-谱-时域变压器(CST-former),通过跨空间、频谱和时间维度的多维注意力机制,增强模型学习事件检测与DoA估计所需领域信息的能力。本文进一步改进CST-former,引入多尺度展开局部嵌入(MSULE)模块,以捕获并聚合多时间-频率尺度上的领域信息。同时,提出针对有限训练数据的微调与后处理策略。通过在STARSS22和STARSS23数据集上的深入消融实验与详尽分析,验证了多维注意力在SELD任务中的有效性。实验证明,所提方法在不使用外部数据的情况下,仍能实现优异性能。

原文摘要 · Abstract (English)

Sound event localization and detection (SELD) is a task for the classification of sound events and the identification of direction of arrival (DoA) utilizing multichannel acoustic signals. For effective classification and localization, a channel-spectro-temporal transformer (CST-former) was suggested. CST-former employs multidimensional attention mechanisms across the spatial, spectral, and temporal domains to enlarge the model's capacity to learn the domain information essential for event detection and DoA estimation over time. In this work, we present an enhanced version of CST-former with multiscale unfolded local embedding (MSULE) developed to capture and aggregate domain information over multiple time-frequency scales. Also, we propose finetuning and post-processing techniques beneficial for conducting the SELD task over limited training datasets. In-depth ablation studies of the proposed architecture and detailed analysis on the proposed modules are carried out to validate the efficacy of multidimensional attentions on the SELD task. Empirical validation through experimentation on STARSS22 and STARSS23 datasets demonstrates the remarkable performance of CST-former and post-processing techniques without using external data.

声音检测多维注意力定位识别音频处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。