ISAC构建可逆稳定音频滤波器组,适配机器学习应用。
ISAC: An Invertible and Stable Auditory Filter Bank with Customizable Kernels for ML Integration
- 基于听觉频率尺度设计非线性中心频点与带宽。
- 滤波器核支持自定义时域长度且可作为可学习卷积核。
- 配套逆滤波器组实现完美重构,适合音视频分析合成。
本文提出ISAC,一种可逆且稳定的感知驱动滤波器组,专为融入机器学习框架而设计。其滤波器的中心频率与带宽遵循非线性听觉频率尺度;滤波器核具有用户自定义的最大时域支撑,可作为可学习的卷积核使用;同时存在对应的逆滤波器组,二者构成完美重构对。ISAC提供强大且易用的音频前端,适用于各类应用场景,包括分析-合成系统。
原文摘要 · Abstract (English)
This paper introduces ISAC, an invertible and stable, perceptually-motivated filter bank that is specifically designed to be integrated into machine learning paradigms. More precisely, the center frequencies and bandwidths of the filters are chosen to follow a non-linear, auditory frequency scale, the filter kernels have user-defined maximum temporal support and may serve as learnable convolutional kernels, and there exists a corresponding filter bank such that both form a perfect reconstruction pair. ISAC provides a powerful and user-friendly audio front-end suitable for any application, including analysis-synthesis schemes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。