arXiv:2410.21561cs.SDcs.LG2024-10ICML被引 2

用定制卷积网络识别低特征音频谱图,提升小样本下的分类精度。

Audio Classification of Low Feature Spectrograms Utilizing Convolutional Neural Networks

  • 设计新型卷积架构,支持二值、单类和相似度学习
  • 在有限或偏斜数据下实现优于传统方法的准确率
  • 适合真实场景中数据稀缺的音频分类任务

当前音频信号分类技术难以处理低特征音频谱图,且依赖于全面但非现实分布的数据集。本文提出若干开创性机器学习方法,用于分析在归一化、偏斜甚至数据量有限条件下的音频谱图。特别地,设计多种定制化卷积神经网络架构,结合二值、单类与孪生学习策略,提取音频信号的谱图特征。实验表明,该方法在分类准确率与计算效率上均达到当前最优水平,显著优于传统音频分类方法。

原文摘要 · Abstract (English)

Modern day audio signal classification techniques lack the ability to classify low feature audio signals in the form of spectrographic temporal frequency data representations. Additionally, currently utilized techniques rely on full diverse data sets that are often not representative of real-world distributions. This paper derives several first-of-its-kind machine learning methodologies to analyze these low feature audio spectrograms given data distributions that may have normalized, skewed, or even limited training sets. In particular, this paper proposes several novel customized convolutional architectures to extract identifying features using binary, one-class, and siamese approaches to identify the spectrographic signature of a given audio signal. Utilizing these novel convolutional architectures as well as the proposed classification methods, these experiments demonstrate state-of-the-art classification accuracy and improved efficiency than traditional audio classification methods.

音频分类卷积网络小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。