利用城市特有声音特征提升声景分类准确率
Improving Acoustic Scene Classification with City Features
- 通过知识蒸馏迁移预训练城市分类模型中的城市特征
- 在三个DCASE数据集上显著提升轻量CNN的分类精度
- 适合关注跨城市声学差异与模型泛化能力的研究者
声景录音通常来自多个不同城市。现有声景分类(ASC)方法主要关注跨城市的通用声学模式以增强泛化能力,却忽略了城市特有的环境与文化因素带来的声学差异。本文假设这些城市特异性特征对ASC任务有益,而非噪声或偏差。为此,提出City2Scene框架,通过知识蒸馏将预训练城市分类模型中的城市知识迁移到场景分类模型中。在DCASE Challenge Task 1的三个包含场景与城市标签的数据集上进行评估,结果表明城市特征为场景分类提供了有价值信息。通过蒸馏城市特异性知识,City2Scene在多种轻量级CNN主干网络上均有效提升性能,达到近年DCASE挑战赛顶尖方案的竞争力。
原文摘要 · Abstract (English)
Acoustic scene recordings are often collected from a diverse range of cities. Most existing acoustic scene classification (ASC) approaches focus on identifying common acoustic scene patterns across cities to enhance generalization. However, the potential acoustic differences introduced by city-specific environmental and cultural factors are overlooked. In this paper, we hypothesize that the city-specific acoustic features are beneficial for the ASC task rather than being treated as noise or bias. To this end, we propose City2Scene, a novel framework that leverages city features to improve ASC. Unlike conventional approaches that may discard or suppress city information, City2Scene transfers the city-specific knowledge from pre-trained city classification models to scene classification model using knowledge distillation. We evaluate City2Scene on three datasets of DCASE Challenge Task 1, which include both scene and city labels. Experimental results demonstrate that city features provide valuable information for classifying scenes. By distilling city-specific knowledge, City2Scene effectively improves accuracy across a variety of lightweight CNN backbones, achieving competitive performance to the top-ranked solutions of DCASE Challenge in recent years.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。