自引导音源分离与分类框架,实现无外部标注的精准目标声音提取。
Self-Guided Target Sound Extraction and Classification Through Universal Sound Separation Model and Multiple Clues
- 通过自循环机制,让模型自主确定目标声音并逐步优化分离与分类结果。
- 在官方数据集上实现CA-SDRi提升11.00 dB,标签准确率达55.8%。
- 适合声音场景分割、自动音频分析等需要端到端精确定位的任务研究者。
本文提出一种多阶段自引导框架,用于应对DCASE 2025 Task 4挑战中的声景空间语义分割(S5)任务。该框架整合了通用声音分离(USS)、单标签分类(SC)和目标声音提取(TSE)三类模型。首先,USS将复杂音频混合信号分解为独立的源波形;每个分离出的波形经由SC模块处理,生成对应的波形与类别标签;这些信息作为TSE的输入,用于提取匹配目标的声源。由于输入信息由系统内部生成,无需外部标注即可自主确定目标。提取出的波形可循环回分类任务,形成迭代优化闭环,持续提升分离质量与分类准确率。因此,该框架被称为多阶段自引导系统。在官方评估数据集上,所提方法在类相关信干比改善(CA-SDRi)上提升11.00 dB,标签预测准确率达55.8%,优于ResUNetK基线4.4 dB和4.3%,并取得所有提交方案中的第一名。
原文摘要 · Abstract (English)
This paper introduces a multi-stage self-directed framework designed to address the spatial semantic segmentation of sound scene (S5) task in the DCASE 2025 Task 4 challenge. This framework integrates models focused on three distinct tasks: Universal Sound Separation (USS), Single-label Classification (SC), and Target Sound Extraction (TSE). Initially, USS breaks down a complex audio mixture into separate source waveforms. Each of these separated waveforms is then processed by a SC block, generating two critical pieces of information: the waveform itself and its corresponding class label. These serve as inputs for the TSE stage, which isolates the source that matches this information. Since these inputs are produced within the system, the extraction target is identified autonomously, removing the necessity for external guidance. The extracted waveform can be looped back into the classification task, creating a cycle of iterative refinement that progressively enhances both separability and labeling accuracy. We thus call our framework a multi-stage self-guided system due to these self-contained characteristics. On the official evaluation dataset, the proposed system achieves an 11.00 dB increase in class-aware signal-to-distortion ratio improvement (CA-SDRi) and a 55.8\% accuracy in label prediction, outperforming the ResUNetK baseline by 4.4 dB and 4.3\%, respectively, and achieving first place among all submissions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。