提出一种统一的2D卷积语音前端,参数高效且适合资源受限场景。
Unified Learnable 2D Convolutional Feature Extraction for ASR
- 用统一的2D卷积结构替代多层混合架构,减少对传统方法的依赖。
- 在有限计算资源下性能接近现有监督学习特征提取器。
- 适合轻量级部署,无需大规模无标签音频预训练。
神经前端为自动语音识别(ASR)系统提供了有前景的特征提取方法,可学习特定任务的特征。然而,现有技术仍受经典方法强烈影响。本文旨在设计一种更通用的前端,统一现有多种来源的层拓扑结构。实验表明,通过减少对已有技术的依赖,可实现通用前端。所提出的2D卷积前端参数高效,适用于计算资源受限场景,无需在大量无标签音频上预训练。结果证明该统一方法不仅可行,且性能可与现有监督学习特征提取器相当。
原文摘要 · Abstract (English)
Neural front-ends represent a promising approach to feature extraction for automatic speech recognition (ASR) systems as they enable to learn specifically tailored features for different tasks. Yet, many of the existing techniques remain heavily influenced by classical methods. While this inductive bias may ease the system design, our work aims to develop a more generic front-end for feature extraction. Furthermore, we seek to unify the front-end architecture contrasting with existing approaches that apply a composition of several layer topologies originating from different sources. The experiments systematically show how to reduce the influence of existing techniques to achieve a generic front-end. The resulting 2D convolutional front-end is parameter-efficient and suitable for a scenario with limited computational resources unlike large models pre-trained on unlabeled audio. The results demonstrate that this generic unified approach is not only feasible but also matches the performance of existing supervised learnable feature extractors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。