为音频事件检测设计高效预训练流程,显著提升模型性能。
Effective Pre-Training of Audio Transformers for Sound Event Detection
- 在AudioSet帧级标注上设计精细化训练流程
- 五种Transformer模型均实现显著性能提升
- 适合音频事件检测研究者直接微调使用
我们提出一种针对音频频谱变换器的预训练流程,用于帧级音频事件检测任务。在常规预训练步骤基础上,引入基于AudioSet帧级标注的精心设计训练方案,包括平衡采样、激进数据增强和集成知识蒸馏。针对五种Transformer模型,我们在AudioSet帧级预测及下游帧级音频事件检测任务上均获得显著优于先前检查点的性能,验证了该流程的有效性。我们发布了由此产生的检查点,研究人员可直接微调以构建高性能音频事件检测模型。
原文摘要 · Abstract (English)
We propose a pre-training pipeline for audio spectrogram transformers for frame-level sound event detection tasks. On top of common pre-training steps, we add a meticulously designed training routine on AudioSet frame-level annotations. This includes a balanced sampler, aggressive data augmentation, and ensemble knowledge distillation. For five transformers, we obtain a substantial performance improvement over previously available checkpoints both on AudioSet frame-level predictions and on frame-level sound event detection downstream tasks, confirming our pipeline's effectiveness. We publish the resulting checkpoints that researchers can directly fine-tune to build high-performance models for sound event detection tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。