arXiv:2512.03637cs.SDcs.LG2025-12中稿 · publication in IEE…被引 1

解决音频谱图变换器中的混叠问题,提升自监督预训练效果

AaSP: Aliasing-aware Self-Supervised Pre-Training for Audio Spectrogram Transformers

  • 引入混叠感知的分块表示,动态融合易混叠频带特征
  • 在多个基准上实现语音、环境音和音乐识别的最优性能
  • 适合需要高鲁棒性音频表征的自监督学习研究者

基于Transformer的音频自监督学习模型通常使用谱图、视觉风格的Transformer和掩码建模目标。然而,时序下采样的卷积分块会降低有效奈奎斯特频率并引入混叠,而简单的低通滤波可能去除任务相关的高频线索。本文提出AaSP,一种面向音频谱图变换器的混叠感知自监督预训练框架。AaSP结合混叠感知分块表示、师生掩码建模、交叉注意力预测器和多掩码对比正则化,学习在掩码视图间保持稳定的表征。其分块嵌入模块AaPE通过带限复正弦核与双侧指数窗,从易混叠调制带中提取特征并融合到标准分块令牌中。核的频率与衰减参数由输入自适应估计,实现可变子带分析。在AudioSet上预训练后,通过微调和线性评估在声学/环境、语音和音乐识别基准上进行测试。微调结果表明,AaSP在AS-20K、ESC-50和NSynth上优于其他自监督基线,其余任务也保持竞争力;线性评估显示在US8K和NSynth上同样有增益。整体而言,AaSP学习到的表征在混叠敏感的时间扰动下更稳定,且在下游迁移任务中表现优异。

原文摘要 · Abstract (English)

Transformer-based audio self-supervised learning (SSL) models commonly use spectrograms, vision-style Transformers, and masked modeling objectives. However, convolutional patchification with temporal downsampling lowers the effective Nyquist frequency and introduces aliasing, while naïve low-pass filtering may remove task-relevant high-frequency cues. We present AaSP, an aliasing-aware self-supervised pre-training framework for audio spectrogram transformers. AaSP combines an aliasing-aware patch representation, teacher-student masked modeling, a cross-attention predictor, and multi-mask contrastive regularization to learn representations that integrate features from alias-prone modulation bands while remaining stable across masked views. Its patch-embedding module, Aliasing-aware Patch Embedding (AaPE), augments standard patch tokens with features from alias-prone modulation bands using a band-limited complex sinusoidal kernel with a two-sided exponential window. The kernel's frequency and decay parameters are estimated from the input, enabling adaptive subband analysis whose outputs are fused with standard patch tokens. We pre-train on AudioSet and evaluate the learned representations by fine-tuning and linear evaluation on acoustic/environmental, speech, and music recognition benchmarks. Under fine-tuning, the full AaSP framework achieves state-of-the-art results on AS-20K, ESC-50, and NSynth among compared self-supervised baselines, while remaining competitive elsewhere. Linear evaluation shows a similar trend, including gains on US8K and NSynth. Overall, AaSP learns representations that are more stable under aliasing-sensitive temporal perturbations and competitive for downstream transfer.

自监督学习音频处理谱图变换器混叠抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。