解析音频自监督学习中目标、结构与应用的匹配关系。
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning

- 按预训练目标分类五种音频自监督范式
- 揭示不同网络结构如何适配特定学习需求
- 适合研究音频模型设计与跨领域泛化者
本文从预训练目标、模型结构归纳偏置与下游应用的对齐角度,重新审视音频自监督学习(SSL)。不同于将SSL视为一系列前文本任务或模型家族的演进,本文关注不同监督信号如何塑造模型应学习的表征。围绕五种范式展开:辅助任务、对比学习、生成重构、离散标记预测与多模态对齐。这些目标对模型提出各异要求,涵盖局部结构敏感性、对比不变性、上下文推断、离散语义抽象及多模态锚定。文章将这些需求与CNN、递归与状态空间模型、Transformer及混合架构的偏置关联,说明局部声学压缩、序列状态传播、内容依赖的全局路由、局部-全局融合分别支持不同形式的音频SSL。该视角也被用于解读语音处理、环境声音分析、音乐信息检索、医疗与生物声学分析及多模态音频理解等下游任务,作为验证表征与架构选择是否具备跨域泛化能力的实际测试。同时回顾基准协议与开放挑战,包括标记化瓶颈、长上下文效率、鲁棒性及安全多模态部署,并探讨基于编码器的标记化与音频-语言建模如何拓展这一目标-架构-应用链路。配套代码库已公开于 https://github.com/colaudiolab/Awesome-Self-Supervised-Audio-Learning。
原文摘要 · Abstract (English)
This paper examines audio self-supervised learning (SSL) through the alignment between pretraining objectives, architectural inductive biases, and downstream applications. Rather than treating SSL methods as a chronological sequence of pretext tasks or model families, we ask how different supervisory signals shape the representations that models are expected to learn. The discussion is organized around five paradigms: auxiliary tasks, contrastive learning, generative reconstruction, discrete token prediction, and multimodal alignment. These objectives place different demands on the model, from local structural sensitivity and contrastive invariance to contextual inference, discrete semantic abstraction, and multimodal grounding. We relate these demands to the biases of CNNs, recurrent and State Space Models, Transformers, and hybrid architectures, showing how local acoustic compression, sequential state propagation, content-dependent global routing, and local--global integration support different forms of audio SSL. The same view is then used to interpret downstream applications in speech processing, environmental sound analysis, music information retrieval, medical and bioacoustic analysis, and multimodal audio understanding as practical tests of whether learned representations and architectural choices generalize across domains. We also review benchmark protocols and open challenges, including tokenization bottlenecks, long-context efficiency, robustness, and secure multimodal deployment, and discuss how codec-based tokenization and audio-language modeling extend this objective--architecture--application pipeline. The accompanying repository is released at https://github.com/colaudiolab/Awesome-Self-Supervised-Audio-Learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。