构建音效分类新标准,提升复杂声音的识别准确率。
Heterogeneous sound classification with the Broad Sound Taxonomy and Dataset
- 采用两级28类音效分类体系,覆盖真实场景下的多样化声音。
- 基于预训练模型的声学语义嵌入比传统特征表现更优,准确率更高。
- 揭示错误根源并提出改进方向,适合音频理解与智能系统研究者参考。
自动声音分类在机器听觉中具有广泛应用,支持情境感知的音频处理与理解。本文研究高类别内差异性声音的自动分类方法,采用包含28个类别的宽泛声音分类体系(Broad Sound Taxonomy),该体系具备语义区分度,适用于实际用户场景。通过人工标注构建数据集,确保类内多样性与现实相关性。对比多种传统与现代机器学习方法,建立该任务基线。重点分析输入特征,比较声学特征与预训练深度神经网络提取的融合声学-语义嵌入的表现。实验表明,融合声学与语义信息的嵌入显著提升分类准确率。通过对分类错误的深入分析,识别失败原因并提出缓解策略。论文强调需深化对分类全链路的理解,重视数据特性,发展能应对复杂性并在真实环境中泛化的有效方法。
原文摘要 · Abstract (English)
Automatic sound classification has a wide range of applications in machine listening, enabling context-aware sound processing and understanding. This paper explores methodologies for automatically classifying heterogeneous sounds characterized by high intra-class variability. Our study evaluates the classification task using the Broad Sound Taxonomy, a two-level taxonomy comprising 28 classes designed to cover a heterogeneous range of sounds with semantic distinctions tailored for practical user applications. We construct a dataset through manual annotation to ensure accuracy, diverse representation within each class and relevance in real-world scenarios. We compare a variety of both traditional and modern machine learning approaches to establish a baseline for the task of heterogeneous sound classification. We investigate the role of input features, specifically examining how acoustically derived sound representations compare to embeddings extracted with pre-trained deep neural networks that capture both acoustic and semantic information about sounds. Experimental results illustrate that audio embeddings encoding acoustic and semantic information achieve higher accuracy in the classification task. After careful analysis of classification errors, we identify some underlying reasons for failure and propose actions to mitigate them. The paper highlights the need for deeper exploration of all stages of classification, understanding the data and adopting methodologies capable of effectively handling data complexity and generalizing in real-world sound environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。