提出新评估指标CASA-SDR,更好区分声音分离与分类错误。
Metric Analysis for Spatial Semantic Segmentation of Sound Scenes
- 先匹配音源再计算分类误差,避免标签混淆干扰评估。
- 在DCASE 2025挑战中验证,相比CA-SDR更准确反映分离性能。
- 适合研究声音场景分割的评估方法,尤其关注分离质量。
空间语义音景分割(S5)需同时完成多通道音频混合中的声源分离与声音事件分类。现有评估方法分别使用分离与分类指标,难以系统对比;而联合指标如类感知信噪比(CA-SDR)会混淆分离与分类错误。尤其当依赖预测标签进行音源匹配时,即使分离效果良好,标签误判也会被掩盖。本文提出类与音源感知信噪比(CASA-SDR),先进行排列不变的音源匹配,再计算分类误差,实现从分类导向到分离导向的转变。我们在理想分离与合成分类错误、源间串扰等受控场景下分析了CA-SDR表现,并与经典SDR和CASA-SDR对比。还引入基于错误与基于音源的聚合策略研究分类误差影响。最后在DCASE 2025挑战任务4的系统上比较了两种指标,发现CA-SDR常过度惩罚标签互换或分离不佳的情况,而CASA-SDR提供更清晰的分离性能评估。
原文摘要 · Abstract (English)
Spatial semantic segmentation of sound scenes (S5) consists of jointly performing audio source separation and sound event classification from a multichannel audio mixture. Evaluating S5 systems with separation and classification metrics individually makes system comparison difficult, whereas existing joint metrics, such as the class-aware signal-to-distortion ratio (CA-SDR), can conflate separation and labeling errors. In particular, CA-SDR relies on predicted class labels for source matching, which may obscure label swaps or misclassifications when the underlying source estimates remain perceptually correct. In this work, we introduce the class and source-aware signal-to-distortion ratio (CASA-SDR), a new metric that performs permutation-invariant source matching before computing classification errors, thereby shifting from a classification-focused approach to a separation-focused approach. We first analyze CA-SDR in controlled scenarios with oracle separation and synthetic classification errors, as well as under controlled cross-contamination between sources, and compare its behavior to that of the classical SDR and CASA-SDR. We also study the impact of classification errors on the metrics by introducing error-based and source-based aggregation strategies. Finally, we compare CA-SDR and CASA-SDR on systems submitted to Task 4 of the DCASE 2025 challenge, highlighting the cases where CA-SDR over-penalizes label swaps or poorly separated sources, while CASA-SDR provides a more interpretable separation-centric assessment of S5 performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。