arXiv:2606.24512eess.AS2026-06

多阶段协同分离与分类,提升声场语义分割精度。

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues

论文配图:A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues
图 1 · 摘自论文原文
  • 分阶段耦合分离与分类模型,逐级优化音源估计。
  • 测试集达15.51 dB的CAPI-SDRi,源准确率78.62%。
  • 适合音频分离与场景理解任务的研究者参考。

本文描述了针对DCASE 2026挑战赛任务4:声场空间语义分割(S5)所提出的系统。我们设计了一个多阶段框架,每阶段均耦合一个分离模型与一个分类模型。第一阶段直接对多通道混合信号进行源分离与分类,其输出作为两个互补线索传递至后续阶段:(i) 音频线索,即分离出的波形,作为低层声学参考;(ii) 类别线索,即以独热向量编码的预测标签。第三阶段在相同框架下复用第二阶段输出,形成迭代自引导精炼过程。此外,我们引入在大规模音频语料上预训练的细粒度帧级音频嵌入作为额外线索,进一步提升分离性能。在测试集上,该系统实现CAPI-SDRi为15.51 dB,混合物准确率为71.09%,源准确率为78.62%;相比基准分别提升7.02 dB、10.38%p和8.22%p。

原文摘要 · Abstract (English)

This report describes the system proposed for the DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes (S5). Specifically, we develop a multi-stage framework in which each stage couples a separation model with a classification model. The first stage performs source separation and classification directly on the multi-channel mixture. Its outputs are then propagated to the following stage as two complementary clues that progressively refine each target estimate: (i) an enrollment clue, the separated waveform itself, serving as a low-level acoustic reference; and (ii) a class clue, the predicted label encoded as a one-hot vector. The third stage reuses the second-stage outputs under the same scheme, forming an iterative self-guided refinement process. In addition, we use a fine-grained frame-level audio embedding from an audio encoder pretrained on a large audio corpus as an additional clue to further improve the audio separation performance. On the test set, the proposed system achieves a CAPI-SDRi of 15.51 dB, a mixture accuracy of 71.09\%, and a source accuracy of 78.62\%; with an improvement of 7.02 dB, 10.38\%p and 8.22\%p compared with the challenge baseline, respectively.

声场分割音频分离多阶段框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。