arXiv:2507.20036cs.SDcs.LG2025-07被引 2

用少量样本优化音频分类,比零样本更准。

Improving Audio Classification by Transitioning from Zero- to Few-Shot

  • 用类内音频嵌入替代噪声文本嵌入
  • 少样本分类显著优于零样本基线
  • 适合小样本音频识别场景

当前先进的音频分类多采用零样本方法,通过比较音频嵌入与描述类别文本的嵌入来实现。这些嵌入通常由对比学习训练的神经网络生成,以对齐音频与文本表示。然而,为某一音频类别选择最优文本描述极具挑战性,尤其当类别包含多种声音时。本文研究了少样本方法,旨在超越零样本方法的分类精度。具体而言,将同类音频嵌入进行分组并处理,以替换固有噪声的文本嵌入。实验结果表明,少样本分类通常优于零样本基线。

原文摘要 · Abstract (English)

State-of-the-art audio classification often employs a zero-shot approach, which involves comparing audio embeddings with embeddings from text describing the respective audio class. These embeddings are usually generated by neural networks trained through contrastive learning to align audio and text representations. Identifying the optimal text description for an audio class is challenging, particularly when the class comprises a wide variety of sounds. This paper examines few-shot methods designed to improve classification accuracy beyond the zero-shot approach. Specifically, audio embeddings are grouped by class and processed to replace the inherently noisy text embeddings. Our results demonstrate that few-shot classification typically outperforms the zero-shot baseline.

音频分类少样本学习对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。