arXiv:2503.16311cs.LGcs.AI2025-03

用结构化噪声生成掩码,让视频音频自监督学习更高效

Structured-Noise Masked Modeling for Video, Audio and Beyond

  • 用噪声过滤生成符合视频音频特性的结构化掩码
  • 在多种模型上均优于随机掩码,提升表示学习效果
  • 无需调参或数据访问,适合多模态自监督研究者

掩码建模已成为强大的自监督学习框架,但现有方法多依赖随机掩码,忽略了不同模态的结构特性。本文提出基于结构化噪声的掩码方法,该方法自然契合视频与音频的空间、时间及频谱特征。通过将白噪声滤波为特定色彩噪声分布,生成保留模态特性的结构化掩码,无需手工设计规则或访问原始数据。实验表明,该方法在标准与先进掩码建模框架中均实现一致性能提升,且无计算开销,验证了模态感知掩码策略对表征学习的重要性。

原文摘要 · Abstract (English)

Masked modeling has emerged as a powerful self-supervised learning framework, but existing methods largely rely on random masking, disregarding the structural properties of different modalities. In this work, we introduce structured noise-based masking, a simple yet effective approach that naturally aligns with the spatial, temporal, and spectral characteristics of video and audio data. By filtering white noise into distinct color noise distributions, we generate structured masks that preserve modality-specific patterns without requiring handcrafted heuristics or access to the data. Our approach improves the performance of masked video and audio modeling frameworks without any computational overhead. Extensive experiments demonstrate that structured noise masking achieves consistent improvement over random masking for standard and advanced masked modeling methods, highlighting the importance of modality-aware masking strategies for representation learning.

自监督学习视频生成音频建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。