提出新指标评估神经网络对采样率变化的鲁棒性
Local Equivariance Error-Based Metrics for Evaluating Sampling-Frequency-Independent Property of Neural Network
- 基于局部等变误差,量化网络对采样率变化的敏感度
- 在音乐源分离任务中,指标与未训练采样率下的性能下降强相关
- 适合关注音频模型泛化能力的研究者和工程师
基于深度神经网络(DNN)的音频处理方法通常仅在单一采样频率(SF)下训练,因此需对未训练的采样率进行重采样处理。然而,近期研究发现,重采样可能降低未训练采样率下的性能。这一问题被忽视,因多数研究仅评估训练采样率下的表现。本文为评估DNN对采样率变化的鲁棒性(即采样频率无关性,SFI)属性,提出三种基于局部等变误差(LEE)的度量指标。LEE用于衡量DNN对输入变换的鲁棒性。通过将信号重采样作为输入变换,扩展LEE以评估音频源分离方法对重采样的鲁棒性。所提指标聚焦于预测时频掩码的网络组件,实验表明其与未训练采样率下的性能下降高度相关。
原文摘要 · Abstract (English)
Audio signal processing methods based on deep neural networks (DNNs) are typically trained only at a single sampling frequency (SF) and therefore require signal resampling to handle untrained SFs. However, recent studies have shown that signal resampling can degrade performance with untrained SFs. This problem has been overlooked because most studies evaluate only the performance at trained SFs. In this paper, to assess the robustness of DNNs to SF changes, which we refer to as the SF-independent (SFI) property, we propose three metrics to quantify the SFI property on the basis of local equivariance error (LEE). LEE measures the robustness of DNNs to input transformations. By using signal resampling as input transformation, we extend LEE to measure the robustness of audio source separation methods to signal resampling. The proposed metrics are constructed to quantify the SFI property in specific network components responsible for predicting time-frequency masks. Experiments on music source separation demonstrated a strong correlation between the proposed metrics and performance degradation at untrained SFs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。