arXiv:2503.22088eess.AScs.SD2025-03中稿 · EUSIPCO2025被引 8

提出音景空间语义分割基线系统,实现声音事件的定位与分离。

Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes

  • 融合音频标签与标签查询源分离模型,分两步完成音景分割。
  • 在第一阶全向声数据上验证,分离效果优于传统方法。
  • 设计新评估指标,同时衡量声音与标签准确性,适合沉浸式音频研究者。

沉浸式通信取得显著进展,尤其得益于沉浸式语音与音频服务编解码器的发布。为推进其落地,DCASE 2025挑战赛提出了空间语义分割音景(S5)任务,聚焦于空间音景中声音事件的检测与分离。本文探索S5任务的解决方法,提出结合音频标签(AT)与标签查询源分离(LSS)模型的基线系统。我们基于ResUNet架构研究两种LSS方法:一是为每个检测到的事件提取单一声音源,二是并发查询多个声音源。由于S5中的每个分离源由其声音事件类别标签标识,我们提出新的类别感知评估指标,可同步评价声音源与标签的性能。在第一阶全向声数据上的实验表明,所提系统有效,且评估指标具有可靠性。

原文摘要 · Abstract (English)

Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the DCASE 2025 Challenge has recently introduced a task for spatial semantic segmentation of sound scenes (S5), which focuses on detecting and separating sound events in spatial sound scenes. In this paper, we explore methods for addressing the S5 task. Specifically, we present baseline S5 systems that combine audio tagging (AT) and label-queried source separation (LSS) models. We investigate two LSS approaches based on the ResUNet architecture: a) extracting a single source for each detected event and b) querying multiple sources concurrently. Since each separated source in S5 is identified by its sound event class label, we propose new class-aware metrics to evaluate both the sound sources and labels simultaneously. Experimental results on first-order ambisonics spatial audio demonstrate the effectiveness of the proposed systems and confirm the efficacy of the metrics.

音景分割空间音频源分离评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。