用物体特征与语义线索提升声源分离与分类效果。
Sound Separation and Classification with Object and Semantic Guidance
- 双路结构融合分离模型特征与预训练分类模型语义,不微调。
- 在DCASE 2025任务4上达11.19 dB CA-SDRi,优于前人11.00 dB。
- 适合需高语义精度的音频理解场景,如智能听觉系统。
空间语义分割任务旨在从多通道信号中分离并分类声音对象。传统方法将大型分类模型与分离模型级联,通过注入分类标签作为下一轮分离的提示。但该方式存在缺陷:小数据集微调会损失分类模型多样性;分离模型特征与预训练分类器输入不匹配;注入的独热标签语义深度不足,易导致错误传播。为此,我们提出双路径分类器(DPC)架构,将分离模型提取的物体特征与预训练分类模型获取的语义表示相结合,无需微调。同时引入语义线索编码器(SCE),增强注入线索的语义深度。系统在DCASE 2025任务4评估集上达到11.19 dB的CA-SDRi,超越先前最优的11.00 dB,验证了融合分离特征与丰富语义线索的有效性。
原文摘要 · Abstract (English)
The spatial semantic segmentation task focuses on separating and classifying sound objects from multichannel signals. To achieve two different goals, conventional methods fine-tune a large classification model cascaded with the separation model and inject classified labels as separation clues for the next iteration step. However, such integration is not ideal, in that fine-tuning over a smaller dataset loses the diversity of large classification models, features from the source separation model are different from the inputs of the pretrained classifier, and injected one-hot class labels lack semantic depth, often leading to error propagation. To resolve these issues, we propose a Dual-Path Classifier (DPC) architecture that combines object features from a source separation model with semantic representations acquired from a pretrained classification model without fine-tuning. We also introduce a Semantic Clue Encoder (SCE) that enriches the semantic depth of injected clues. Our system achieves a state-of-the-art 11.19 dB CA-SDRi and enhanced semantic fidelity on the DCASE 2025 task4 evaluation set, surpassing the top-rank performance of 11.00 dB. These results highlight the effectiveness of integrating separator-derived features and rich semantic clues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。