arXiv:2410.05302eess.AScs.LG2024-10中稿 · MLSP 2024被引 2

用新方法微调原型网络,提升音频少样本分类效果

Episodic fine-tuning prototypical networks for optimization-based few-shot learning: Application to audio classification

  • 在测试时仅用支持集微调原型网络,不依赖查询集
  • 结合MAML和Meta-Curvature,使模型在少量样本下快速适应
  • 在ESC-50和Speech Commands v2上显著优于传统原型网络

原型网络(ProtoNet)因其优异性能和简单实现,已成为少样本学习(FSL)中的主流方法。本文提出一种新颖的微调策略:在测试任务的支撑集上对ProtoNet进行微调,不使用仅用于评估的查询集。进一步设计了将ProtoNet与基于优化的少样本算法(如MAML和Meta-Curvature)结合的框架,利用这些算法赋予模型从极少样本中快速适应的能力。通过引入专门的阶段性微调机制,以原型网络为目标模型,显著提升了其微调表现。实验表明,在ESC-50和Speech Commands v2数据集上的音频少样本分类任务中,所提方法MAML-Proto与MC-Proto均显著优于标准ProtoNet。尽管当前应用仅限于音频领域,但该方法具有通用性,可轻松拓展至其他任务域。

原文摘要 · Abstract (English)

The Prototypical Network (ProtoNet) has emerged as a popular choice in Few-shot Learning (FSL) scenarios due to its remarkable performance and straightforward implementation. Building upon such success, we first propose a simple (yet novel) method to fine-tune a ProtoNet on the (labeled) support set of the test episode of a C-way-K-shot test episode (without using the query set which is only used for evaluation). We then propose an algorithmic framework that combines ProtoNet with optimization-based FSL algorithms (MAML and Meta-Curvature) to work with such a fine-tuning method. Since optimization-based algorithms endow the target learner model with the ability to fast adaption to only a few samples, we utilize ProtoNet as the target model to enhance its fine-tuning performance with the help of a specifically designed episodic fine-tuning strategy. The experimental results confirm that our proposed models, MAML-Proto and MC-Proto, combined with our unique fine-tuning method, outperform regular ProtoNet by a large margin in few-shot audio classification tasks on the ESC-50 and Speech Commands v2 datasets. We note that although we have only applied our model to the audio domain, it is a general method and can be easily extended to other domains.

少样本学习音频分类原型网络微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。