一个可弹性伸缩的语音感知模型,支持多语言多场景识别。
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
- 采用弹性专家混合架构,一次训练即可按需缩放部署规模。
- 利用自监督数据生成,训练数据达数百万小时,使错误率从4.98%降至2.45%。
- 不仅识别中文语音,还能感知语种、方言、情绪、性别和声音事件。
大型自动语音识别(ASR)模型需要大量参数、海量数据和高算力训练,仅能部署于高性能云平台,功能局限且成本高昂。本文首次提出弹性专家混合(eMoE)模型,只需一次训练即可根据部署需求弹性缩放。同时,设计无监督数据生成与验证流程,从多个领域收集数百万小时音频用于训练。结合两项技术,系统实现弹性部署能力,并将SpeechIO测试集上的字符错误率(CER)从4.98%降至2.45%。此外,该模型不仅能完成普通话识别,还具备多语言、多方言、情绪、性别及声音事件感知能力,我们将其称为自动语音感知(ASP),实验部分展示了相关感知结果。
原文摘要 · Abstract (English)
Large Automatic Speech Recognition (ASR) models demand a vast number of parameters, copious amounts of data, and significant computational resources during the training process. However, such models can merely be deployed on high-compute cloud platforms and are only capable of performing speech recognition tasks. This leads to high costs and restricted capabilities. In this report, we initially propose the elastic mixture of the expert (eMoE) model. This model can be trained just once and then be elastically scaled in accordance with deployment requirements. Secondly, we devise an unsupervised data creation and validation procedure and gather millions of hours of audio data from diverse domains for training. Using these two techniques, our system achieves elastic deployment capabilities while reducing the Character Error Rate (CER) on the SpeechIO testsets from 4.98\% to 2.45\%. Thirdly, our model is not only competent in Mandarin speech recognition but also proficient in multilingual, multi-dialect, emotion, gender, and sound event perception. We refer to this as Automatic Speech Perception (ASP), and the perception results are presented in the experimental section.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。