arXiv:2604.00776eess.AS2026-04被引 2

解决复杂声场中声音事件的定位与分离,提升沉浸式音频体验。

Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes

  • 联合检测与分离多源声音事件,支持同类型多个声源。
  • 引入无目标声源和多同类型声源场景,更贴近真实环境。
  • 提供完整数据集与评测指标,适合音频感知与空间建模研究者。

本文介绍了 DCASE 2026 挑战赛任务 4——空间语义声音场景分割(S5)的设置。该任务聚焦于复杂空间音频混合中的声音事件联合检测与分离,为沉浸式通信奠定基础。自 DCASE 2025 首次提出后,2026 年的任务在设置上进行了关键更新:允许混合音频中包含多个同类别声源,且可不包含目标声源,以更好地反映真实应用场景。本文详细说明了任务设定、评估指标调整及数据集更新,并报告与分析了参赛系统的实验结果。官方数据与代码获取地址为 https://github.com/nttcslab/dcase2026_task4_baseline。

原文摘要 · Abstract (English)

This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 4, Spatial Semantic Segmentation of Sound Scenes (S5). The S5 task focuses on the joint detection and separation of sound events in complex spatial audio mixtures, contributing to the foundation of immersive communication. First introduced in DCASE 2025, the S5 task continues in DCASE 2026 Task 4 with key changes to better reflect real-world conditions, including allowing mixtures to contain multiple sources of the same class and to contain no target sources. In this paper, we describe task setting, along with the corresponding updates to the evaluation metrics and dataset. The experimental results of the submitted systems are also reported and analyzed. The official access point for data and code is https://github.com/nttcslab/dcase2026_task4_baseline.

声音分割空间音频声学事件DCASE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。