arXiv:2506.10676cs.SDeess.AS2025-06被引 10

解决多通道声音场景的空间语义分割问题,实现声音事件的精准定位与分离。

Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes

  • 基于多通道信号,将混合声音事件按空间位置和类型进行分离。
  • 构建了新数据集并验证系统在真实场景下的性能表现。
  • 适合研究声音感知、智能音频处理与沉浸式通信的学者参考。

空间语义分割声音场景(S5)旨在提升从多通道输入信号中检测与分离混合声音事件的能力,同时保留其空间信息,为沉浸式通信提供基础支持。最终目标是将具有6自由度(6DoF)信息的声音事件信号分解为干声对象信号及其元数据,包括事件类别和方向等空间信息。由于已有挑战任务覆盖部分功能,本年度任务聚焦于从多通道空间输入信号中检测并分离声音事件。本文介绍了DCASE 2025挑战赛第4任务的S5任务设定及新构建的DCASE2025 Task 4数据集,该数据集为本任务专门录制与整理。同时报告了在该数据集上训练与评估的S5系统的实验结果。完整论文将在挑战结果公布后正式发表。

原文摘要 · Abstract (English)

Spatial Semantic Segmentation of Sound Scenes (S5) aims to enhance technologies for sound event detection and separation from multi-channel input signals that mix multiple sound events with spatial information. This is a fundamental basis of immersive communication. The ultimate goal is to separate sound event signals with 6 Degrees of Freedom (6DoF) information into dry sound object signals and metadata about the object type (sound event class) and representing spatial information, including direction. However, because several existing challenge tasks already provide some of the subset functions, this task for this year focuses on detecting and separating sound events from multi-channel spatial input signals. This paper outlines the S5 task setting of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2025 Challenge Task 4 and the DCASE2025 Task 4 Dataset, newly recorded and curated for this task. We also report experimental results for an S5 system trained and evaluated on this dataset. The full version of this paper will be published after the challenge results are made public.

声音分割空间感知多通道处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。