用自编码器压缩语音模型特征,提速8倍还更准
Autoencoder based optimized SSL representations: Complexity Minimization and improved Dysarthric ASR

- 用自编码器压缩高维语音特征,降低计算复杂度
- 在构音障碍语音识别中,词错误率更低,训练快8倍
- 适合算力有限的语音识别场景,如边缘设备
自监督学习(SSL)模型虽能提取丰富语音表征,但特征维度高,带来高计算开销。本文提出一种基于自编码器的SSL-AE瓶颈方法,将高维SSL特征映射到紧凑空间,在保持构音障碍语音识别(ASR)性能的同时显著降低计算复杂度和训练时间。实验表明,该方法相较原始SSL基线可将训练时间缩短8倍,同时维持甚至降低词错误率(WER),验证了自编码器在资源受限环境下对SSL特征压缩的有效性。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) models extract rich speech representations but often come with high-dimensional features, increasing computational complexity. This work explores an SSL-AutoEncoder (SSL-AE) bottlenecking approach to efficiently reduce feature dimensions while maintaining dysarthric Automatic Speech Recognition (ASR) performance. By leveraging an autoencoder, we transform high-dimensional SSL features into a compact space, reducing model complexity and training time. Our method preserves essential speech information, achieving reduced Word Error Rates (WER) while significantly lowering computational costs. Experiments show SSL-AE bottlenecking reduces training time by 8x compared to the SSL baseline, demonstrating efficiency without sacrificing recognition performance. These results highlight AE as an effective solution for SSL feature compression in resource-constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。