将人耳听觉模型可微化,实现高效且可解释的音频处理。
Biomimetic Frontend for Differentiable Audio Processing
- 基于人耳听觉机制构建可微分前端,融合传统信号处理与深度学习
- 仅用少量数据即可训练,在分类与增强任务中表现优于黑箱模型
- 适合追求可解释性与低资源训练的音频应用开发者
尽管音频与语音处理模型日益深入和端到端化,但随之带来的是高昂的训练成本和脆弱性。本文基于经典的人类听觉模型,将其改造为可微分形式,从而将传统可解释的仿生信号处理方法与深度学习框架相结合。该方法使模型在较小数据量下即可高效训练,兼具表达力与可解释性。我们在音频分类与增强等任务上验证了该模型,结果表明其在计算效率和鲁棒性方面超越黑箱方法,即使在少量训练数据条件下也表现优异。此外,本文还探讨了其他潜在应用场景。
原文摘要 · Abstract (English)
While models in audio and speech processing are becoming deeper and more end-to-end, they as a consequence need expensive training on large data, and are often brittle. We build on a classical model of human hearing and make it differentiable, so that we can combine traditional explainable biomimetic signal processing approaches with deep-learning frameworks. This allows us to arrive at an expressive and explainable model that is easily trained on modest amounts of data. We apply this model to audio processing tasks, including classification and enhancement. Results show that our differentiable model surpasses black-box approaches in terms of computational efficiency and robustness, even with little training data. We also discuss other potential applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。