提出轻量级频谱稀疏掩码,提升音频自监督学习性能。
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
- 基于音频频谱稀疏性设计加权掩码,降低计算开销。
- 相比传统块掩码,新方法在事件理解任务上表现更优。
- 适合追求高效高精度音频表征的开发者使用。
自掩码自动编码器提出以来,多种掩码技术不断改进。本文重新思考基于掩码预测的音频自监督学习(SSL)中掩码策略的设计,针对通用音频频谱图进行研究。尽管近期的智能掩码方法受到关注,但其带来显著的计算开销。受此启发,本文提出一种轻量级的分散加权掩码(DWM),利用音频内容固有的频谱稀疏特性。实验表明,反向块掩码(inverse block masking)虽能提升音频事件理解性能,却牺牲了泛化能力。DWM有效缓解了这些局限与计算复杂度,实现一致的性能提升。本工作为基于掩码预测的音频表征学习中的掩码策略设计提供了实用指导。
原文摘要 · Abstract (English)
Since the introduction of Masked Autoencoders, various improvements to masking techniques have been explored. In this paper, we rethink masking strategies for audio representation learning using masked prediction-based self-supervised learning (SSL) on general audio spectrograms. While recent informed masking techniques have attracted attention, we observe that they incur substantial computational overhead. Motivated by this observation, we propose dispersion-weighted masking (DWM), a lightweight masking strategy that leverages the spectral sparsity inherent in the frequency structure of audio content. Our experiments show that inverse block masking, commonly used in recent SSL frameworks, improves audio event understanding performance while introducing a trade-off in generalization. The proposed DWM alleviates these limitations and computational complexity, leading to consistent performance improvements. This work provides practical guidance on masking strategy design for masked prediction-based audio representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。