用语义层级引导半监督学习,提升环境音分类准确率
ECHO: Environmental Sound Classification with Hierarchical Ontology-guided Semi-Supervised Learning
- 基于标签本体构建层次化预训练任务,利用大模型生成粗粒度标签
- 在三个数据集上相比基线提升1%至8%的分类准确率
- 适合关注少标注环境下音频分类的研究者
环境声音分类是信号处理领域的经典问题,以往研究多集中于全监督方法。近年来,半监督与自监督方法兴起,但普遍依赖大量未标注数据才能提升性能。本文提出一种新框架ECHO(Environmental Sound Classification with Hierarchical Ontology-guided semi-Supervised Learning),利用标签本体的层次结构设计新型预训练任务:通过大语言模型(LLM)根据真实标签本体生成粗粒度标签,让模型学习语义表示。训练后的模型再经有监督微调以完成实际分类任务。该方法在UrbanSound8K、ESC-10和ESC-50三个数据集上均取得1%至8%的准确率提升。
原文摘要 · Abstract (English)
Environment Sound Classification has been a well-studied research problem in the field of signal processing and up till now more focus has been laid on fully supervised approaches. Over the last few years, focus has moved towards semi-supervised methods which concentrate on the utilization of unlabeled data, and self-supervised methods which learn the intermediate representation through pretext task or contrastive learning. However, both approaches require a vast amount of unlabelled data to improve performance. In this work, we propose a novel framework called Environmental Sound Classification with Hierarchical Ontology-guided semi-supervised Learning (ECHO) that utilizes label ontology-based hierarchy to learn semantic representation by defining a novel pretext task. In the pretext task, the model tries to predict coarse labels defined by the Large Language Model (LLM) based on ground truth label ontology. The trained model is further fine-tuned in a supervised way to predict the actual task. Our proposed novel semi-supervised framework achieves an accuracy improvement in the range of 1\% to 8\% over baseline systems across three datasets namely UrbanSound8K, ESC-10, and ESC-50.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。