arXiv:2510.18206eess.AScs.SD2025-10中稿 · ICASSP2026被引 1

动态调整音频前端,让模型在复杂环境更稳定。

Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing

  • 用神经控制器实时调节能量归一化参数,实现推理时自适应。
  • 在多个音频分类任务中,复杂环境下性能优于固定与可学习前端。
  • 适合需要高鲁棒性的语音识别、声学检测等实际应用。

在音频信号处理中,可学习的前端已在多种任务中展现出强大性能,但其参数在训练完成后固定不变,缺乏推理时的灵活性,限制了在动态复杂声学环境下的鲁棒性。本文提出一种新型自适应音频前端范式,将静态参数化替换为闭环神经控制器。具体而言,我们简化了可学习前端LEAF架构,并通过动态调节每通道能量归一化(Per-Channel Energy Normalization)引入自适应表示。神经控制器利用当前及缓存的历史子带能量,实现在推理阶段的输入依赖性自适应。在多个音频分类任务上的实验表明,所提出的自适应前端在干净和复杂声学条件下均持续优于以往的固定与可学习前端。结果表明,神经自适应是下一代音频前端的有前景方向。

原文摘要 · Abstract (English)

In audio signal processing, learnable front-ends have shown strong performance across diverse tasks by optimizing task-specific representation. However, their parameters remain fixed once trained, lacking flexibility during inference and limiting robustness under dynamic complex acoustic environments. In this paper, we introduce a novel adaptive paradigm for audio front-ends that replaces static parameterization with a closed-loop neural controller. Specifically, we simplify the learnable front-end LEAF architecture and integrate a neural controller for adaptive representation via dynamically tuning Per-Channel Energy Normalization. The neural controller leverages both the current and the buffered past subband energies to enable input-dependent adaptation during inference. Experimental results on multiple audio classification tasks demonstrate that the proposed adaptive front-end consistently outperforms prior fixed and learnable front-ends under both clean and complex acoustic conditions. These results highlight neural adaptability as a promising direction for the next generation of audio front-ends.

音频处理自适应神经控制器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。