arXiv:2603.25476cs.LG2026-03

研究音频迁移学习中数据量与类别相似性的影响。

How Class Ontology and Data Scale Affect Audio Transfer Learning

  • 在基于语义层级的AudioSet子集上预训练模型
  • 更多样本和类别提升迁移效果,但任务相似性更重要
  • 适合关注音频模型泛化能力的研究者

迁移学习是深度学习中的核心概念,使神经网络能在小数据任务中受益于大规模预训练数据。尽管应用广泛且效果显著,其内在机制仍存在诸多未解问题,尤其对何时及如何有效工作尚不明确。为此,本文针对音频到音频的迁移学习开展严谨研究:在(基于语义层级的)AudioSet子集上预训练多种模型状态,并在三个计算机听觉任务上微调——声学场景识别、鸟类活动识别和语音命令识别。结果表明,增加预训练数据的样本数和类别数均对迁移学习有正向影响,但这种提升通常被预训练与下游任务间的相似性所超越,后者促使模型学习到可比特征。

原文摘要 · Abstract (English)

Transfer learning is a crucial concept within deep learning that allows artificial neural networks to benefit from a large pre-training data basis when confronted with a task of limited data. Despite its ubiquitous use and clear benefits, there are still many open questions regarding the inner workings of transfer learning and, in particular, regarding the understanding of when and how well it works. To that extent, we perform a rigorous study focusing on audio-to-audio transfer learning, in which we pre-train various model states on (ontology-based) subsets of AudioSet and fine-tune them on three computer audition tasks, namely acoustic scene recognition, bird activity recognition, and speech command recognition. We report that increasing the number of samples and classes in the pre-training data both have a positive impact on transfer learning. This is, however, generally surpassed by similarity between pre-training and the downstream task, which can lead the model to learn comparable features.

音频迁移深度学习特征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。