arXiv:2509.15703cs.SDeess.AS2025-09中稿 · ICASSP 2026被引 1

SONAR让音频模型持续学习新领域数据,不丢旧知识。

SONAR: Self-Distilled Continual Pre-training for Domain Adaptive Audio Representation

  • 用联合采样+动态词表扩展,持续学新声音
  • 在4个不同音频领域上既适应新数据又防遗忘
  • 适合长期更新的语音识别与分类系统

大规模数据集(如AudioSet)上的自监督学习已成为音频表征学习的主流范式。尽管不断涌现的新未标注音频为丰富静态表征提供了机会,但直接从头训练模型成本过高,且会丢弃已有模型权重中的宝贵知识。为此,我们提出SONAR(基于BEATs的自蒸馏持续预训练框架),有效适应新领域并缓解灾难性遗忘,解决三大挑战:对新旧数据实施联合采样策略、应用正则化平衡特异性与通用性、动态扩展分词器代码本以捕捉新声学模式。在四个不同领域的实验表明,该方法兼具高适应性与强抗遗忘能力。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) on large-scale datasets like AudioSet has become the dominant paradigm for audio representation learning. While the continuous influx of new, unlabeled audio presents an opportunity to enrich these static representations, a naive approach is to retrain the model from scratch using all available data. However, this method is computationally prohibitive and discards the valuable knowledge embedded in the previously trained model weights. To address this inefficiency, we propose SONAR (Self-distilled cONtinual pre-training for domain adaptive Audio Representation), a continual pre-training framework built upon BEATs. SONAR effectively adapts to new domains while mitigating catastrophic forgetting by tackling three key challenges: implementing a joint sampling strategy for new and prior data, applying regularization to balance specificity and generality, and dynamically expanding the tokenizer codebook for novel acoustic patterns. Experiments across four distinct domains demonstrate that our method achieves both high adaptability and robust resistance to forgetting.

音频表征持续学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。