提出新型探测方法,让音频自监督模型评估更可靠,性能突破新纪录。
BAT: Better Audio Transformer Guided by Convex Gated Probing
- 基于原型的凸门控探测,高效利用冻结层并定位关键特征位置。
- 在AudioSet上实现与微调相当的性能,排名稳定性显著提升。
- 用于优化现有模型,推动音频自监督学习迈向可复现新标准。
探测法在计算机视觉中被广泛用于真实评估自监督学习(SSL)嵌入质量,因为微调可能扭曲其真实表现。相比之下,音频SSL模型仍依赖微调,因简单探测无法释放其全部潜力,且在AudioSet上会改变模型排名。因此,亟需一种稳健高效的探测机制来指导音频SSL的发展方向。本文提出凸门控探测(CGP),一种基于原型的方法,显著缩小了探测与微调之间的差距。CGP通过门控机制高效利用所有冻结层,并揭示潜在任务相关信息的位置。以CGP作为可靠的后验评估探针,我们重构了当前最佳音频模型的整个SSL流程,改进数据预处理、模型架构和预训练方案,提出更好的音频变换器(BAT),并在多个音频基准上建立新最优性能。
原文摘要 · Abstract (English)
Probing is widely adopted in computer vision to faithfully evaluate self-supervised learning (SSL) embeddings, as finetuning may misrepresent their inherent quality. In contrast, audio SSL models still rely on finetuning because simple probing fails to unlock their full potential and alters their rankings when competing on AudioSet. Hence, a robust and efficient probing mechanism is required to guide the trajectory of audio SSL towards reliable and reproducible methods. We introduce Convex Gated Probing (CGP), a prototype-based method that significantly closes the gap between finetuning and probing in audio. CGP efficiently utilizes all frozen layers via a gating mechanism and exposes the location of latent task-relevant information. Guided by CGP as a reliable post-hoc evaluation probe, we rework the entire SSL pipeline of current best performing audio models that use legacy implementations of prior SSL methods. By refining data preprocessing, model architecture, and pretraining recipe, we introduce Better Audio Transformer (BAT), and establish new SOTA on audio benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。