arXiv:2607.16688eess.AS2026-07中稿 · IWAENC 2026被引 2

让音频模型学会识别噪声,提升嘈杂环境下的表现

NABEATs: Noise-Aware Audio Representation Learning

论文配图:NABEATs: Noise-Aware Audio Representation Learning
图 1 · 摘自论文原文
  • 引入参考噪声输入,让模型在推理时感知具体噪声特征
  • 在多种噪声环境下,下游任务性能显著优于传统模型
  • 能有效泛化到未见过的噪声类型,适合真实场景应用

我们提出噪声感知音频自监督学习(Noise-Aware Audio SSL)的概念,目标是在编码音频混合信号的同时抑制不期望的噪声,并提出基于BEATs框架的实现——噪声感知BEATs(NABEATs)。现有音频自监督模型虽能处理广泛音频信号,但在噪声环境下难以聚焦与下游任务相关的有效声音,导致性能下降。为解决此问题,NABEATs通过一个辅助的参考噪声输入,从含噪音频中估计出干净的BEATs表征。该参考噪声使模型能在推理时考虑特定噪声特性,从而在不同运行环境中实现更优泛化能力。实验表明,NABEATs在多种下游任务中显著提升了噪声条件下的性能,并能良好泛化至未见噪声类型。

原文摘要 · Abstract (English)

We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs (NABEATs) as a BEATs-based realization of this framework. Audio SSL models are designed to handle a wide range of audio signals. Consequently, under noisy conditions, they cannot effectively focus on the target sounds relevant to a downstream task, resulting in degraded performance. To address this issue, NABEATs is trained to estimate clean BEATs representations from a noisy audio signal with an auxiliary reference noise input. This reference noise enables the model to account for specific noise characteristics at inference time, thereby achieving better generalization across operating environments. Our experimental evaluations demonstrate that NABEATs significantly improves performance of various downstream tasks under noisy conditions and also generalizes well to unseen noise types.

音频处理自监督学习噪声鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。