让音频前端动态适应环境,比固定结构更稳定可靠
Should Audio Front-ends be Adaptive? Comparing Learnable and Adaptive Front-ends
- 用神经反馈控制器动态调整频谱滤波器的品质因数
- 在多种任务上优于传统可学习前端,且训练过程更稳定
- 适合对环境鲁棒性要求高的语音与声音识别场景
传统音频处理多采用手工设计的特征(如梅尔滤波器组),近年来可学习前端从原始波形直接提取表示逐渐兴起。然而,手工滤波器和现有可学习前端在推理时均采用固定计算图,无法像人耳一样动态适应不同声学环境。为此,本文探讨音频前端是否应具备自适应能力,对比了Ada-FE(一种采用神经自适应反馈控制器动态调节频谱分解滤波器Q值的新型自适应前端)与主流可学习前端。我们在两个常用后端骨干网络和涵盖语音、声音事件、音乐的广泛音频基准上系统评估。结果表明,Ada-FE性能超越先进可学习前端,更重要的是,在不同训练周期下的测试样本上表现出显著的稳定性与鲁棒性。
原文摘要 · Abstract (English)
Hand-crafted features, such as Mel-filterbanks, have traditionally been the choice for many audio processing applications. Recently, there has been a growing interest in learnable front-ends that extract representations directly from the raw audio waveform. \textcolor{black}{However, both hand-crafted filterbanks and current learnable front-ends lead to fixed computation graphs at inference time, failing to dynamically adapt to varying acoustic environments, a key feature of human auditory systems.} To this end, we explore the question of whether audio front-ends should be adaptive by comparing the Ada-FE front-end (a recently developed adaptive front-end that employs a neural adaptive feedback controller to dynamically adjust the Q-factors of its spectral decomposition filters) to established learnable front-ends. Specifically, we systematically investigate learnable front-ends and Ada-FE across two commonly used back-end backbones and a wide range of audio benchmarks including speech, sound event, and music. The comprehensive results show that our Ada-FE outperforms advanced learnable front-ends, and more importantly, it exhibits impressive stability or robustness on test samples over various training epochs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。