arXiv:2410.23008cs.SDeess.AS2024-10中稿 · IEEE ICASSP 2025

从现有音频数据中自动发现新类别,提升下游分类效果。

SoundCollage: Automated Discovery of New Classes in Audio Datasets

  • 通过声音分解与自动化标注,挖掘音频数据中的隐藏类别
  • 新发现类别的下游分类准确率提升最高达34.7%
  • 适合想复用现有数据、拓展音频应用的研究者

开发新机器学习应用常需收集新数据集,但现有数据集可能已包含可用信息。我们提出 SoundCollage 框架,通过(1)音频预处理管道分解音频样本中的不同声音,(2)基于模型的自动化标注机制识别新发现类别。同时引入清晰度度量评估所发现类别的连贯性,以支持下游任务训练。实验表明,新类别样本上的下游音频分类器准确率相比基线最高提升34.7%,在保留数据集上提升4.5%。结果表明,SoundCollage 具有使数据集可重用的潜力,可通过标注新类别实现。为推动该方向研究,代码已开源:https://github.com/nokia-bell-labs/audio-class-discovery。

原文摘要 · Abstract (English)

Developing new machine learning applications often requires the collection of new datasets. However, existing datasets may already contain relevant information to train models for new purposes. We propose SoundCollage: a framework to discover new classes within audio datasets by incorporating (1) an audio pre-processing pipeline to decompose different sounds in audio samples, and (2) an automated model-based annotation mechanism to identify the discovered classes. Furthermore, we introduce the clarity measure to assess the coherence of the discovered classes for better training new downstream applications. Our evaluations show that the accuracy of downstream audio classifiers within discovered class samples and a held-out dataset improves over the baseline by up to 34.7% and 4.5%, respectively. These results highlight the potential of SoundCollage in making datasets reusable by labeling with newly discovered classes. To encourage further research in this area, we open-source our code at https://github.com/nokia-bell-labs/audio-class-discovery.

音频发现自动标注数据复用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。