面向新加坡及东南亚的语音基础模型,支持本地口语与英文语音识别。
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
- 从头预训练20万小时无标注语音数据,采用自监督掩码建模。
- 在新加坡口语和自然语流识别任务中表现优于现有方法。
- 适合研究多语言语音处理、区域语音识别的学者与开发者。
本技术报告介绍了MERaLiON-SpeechEncoder,一个为支持多种下游语音应用而设计的基础模型。作为新加坡国家多模态大语言模型计划的一部分,该模型专为新加坡及周边东南亚地区语音处理需求量身定制。当前主要支持英语,包括新加坡式英语。我们正积极扩展数据集,后续版本将逐步覆盖其他语言。MERaLiON-SpeechEncoder基于自监督学习,从零开始在20万小时未标注语音数据上进行预训练,采用掩码语言建模方法。文中详细描述了训练过程与超参数调优实验。评估结果显示,该模型在自发语流和新加坡口语识别基准上取得显著提升,同时在其他十个语音任务中保持与现有最先进语音编码器相当的性能。我们承诺将公开发布该模型,以推动新加坡及更广泛地区的科研发展。
原文摘要 · Abstract (English)
This technical report describes the MERaLiON-SpeechEncoder, a foundation model designed to support a wide range of downstream speech applications. Developed as part of Singapore's National Multimodal Large Language Model Programme, the MERaLiON-SpeechEncoder is tailored to address the speech processing needs in Singapore and the surrounding Southeast Asian region. The model currently supports mainly English, including the variety spoken in Singapore. We are actively expanding our datasets to gradually cover other languages in subsequent releases. The MERaLiON-SpeechEncoder was pre-trained from scratch on 200,000 hours of unlabelled speech data using a self-supervised learning approach based on masked language modelling. We describe our training procedure and hyperparameter tuning experiments in detail below. Our evaluation demonstrates improvements to spontaneous and Singapore speech benchmarks for speech recognition, while remaining competitive to other state-of-the-art speech encoders across ten other speech tasks. We commit to releasing our model, supporting broader research endeavours, both in Singapore and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。