低采样率音频分离性能下降?新方法让模型更抗干扰。
Dissecting Performance Degradation in Audio Source Separation under Sampling Frequency Mismatch
- 用噪声扰动插值核,模拟高频信息增强
- 在低采样率下性能提升,优于传统重采样
- 无需训练,适合各类音频分离模型使用
基于深度神经网络的音频处理方法通常在单一采样频率(SF)下训练。为应对未训练的采样频率,常采用信号重采样,但当输入采样频率低于训练频率时,性能会显著下降。本文通过两个假设探究原因:(i) 上采样导致高频成分缺失;(ii) 高频存在比精确表示更重要。为此,比较了传统重采样与三种替代方案:重采样后加高斯噪声、噪声核重采样(用高斯噪声扰动插值核以丰富高频成分)、可训练核重采样(通过训练自适应插值核)。音乐源分离实验表明,噪声核与可训练核重采样能有效缓解传统方法的性能退化。进一步证明,噪声核重采样在多种模型中均有效,是一种简单且实用的解决方案。
原文摘要 · Abstract (English)
Audio processing methods based on deep neural networks are typically trained at a single sampling frequency (SF). To handle untrained SFs, signal resampling is commonly employed, but it can degrade performance, particularly when the input SF is lower than the trained SF. This paper investigates the causes of this degradation through two hypotheses: (i) the lack of high-frequency components introduced by up-sampling, and (ii) the greater importance of their presence than their precise representation. To examine these hypotheses, we compare conventional resampling with three alternatives: post-resampling noise addition, which adds Gaussian noise to the resampled signal; noisy-kernel resampling, which perturbs the kernel with Gaussian noise to enrich high-frequency components; and trainable-kernel resampling, which adapts the interpolation kernel through training. Experiments on music source separation show that noisy-kernel and trainable-kernel resampling alleviate the degradation observed with conventional resampling. We further demonstrate that noisy-kernel resampling is effective across diverse models, highlighting it as a simple yet practical option.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。