arXiv:2601.22504eess.AS2026-01中稿 · ICASSP 2026被引 3

解决声音场景中同类别声源导致的分离性能下降问题

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources

  • 提出类感知排列不变损失,让模型处理同类别重复标签查询
  • 新评估指标在含同类别源的混合音频上表现更稳定
  • 适用于真实复杂声学环境下的音景语义分割系统

为推动沉浸式通信发展,DCASE 2025挑战赛引入了空间语义音景分割(S5)任务。该任务将多通道音频混合输入,输出对应类别的单通道干信号。尽管挑战赛限定每组混合音频中类别互斥,但现实场景中常存在多个同类别声源。此类重复标签会显著降低依赖标签查询的源分离(LQSS)模型性能,并影响官方评估指标的有效性。为此,本文提出一种类感知排列不变损失函数,使LQSS模型可处理含重复标签的查询;同时重新设计S5评估指标,消除同类别源带来的歧义。为验证方法有效性,扩展标签预测模型以支持同类别标签。实验表明,所提方法在含与不含同类别源的混合音频上均具优异效果与鲁棒性。

原文摘要 · Abstract (English)

To advance immersive communication, the Detection and Classification of Acoustic Scenes and Events (DCASE) 2025 Challenge recently introduced Task 4 on Spatial Semantic Segmentation of Sound Scenes (S5). An S5 system takes a multi-channel audio mixture as input and outputs single-channel dry sources along with their corresponding class labels. Although the DCASE 2025 Challenge simplifies the task by constraining class labels in each mixture to be mutually exclusive, real-world mixtures frequently contain multiple sources from the same class. The presence of duplicated labels can significantly degrade the performance of the label-queried source separation (LQSS) model, which is the key component of many existing S5 systems, and can also limit the validity of the official evaluation metric of DCASE 2025 Task 4. To address these issues, we propose a class-aware permutation-invariant loss function that enables the LQSS model to handle queries involving duplicated labels. In addition, we redesign the S5 evaluation metric to eliminate ambiguities caused by these same-class sources. To evaluate the proposed method within the S5 system, we extend the label prediction model to support same-class labels. Experimental results demonstrate the effectiveness of the proposed methods and the robustness of the new metric on mixtures both with and without same-class sources.

音景分割源分离评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。