arXiv:2501.01034cs.CLcs.SD2025-01被引 13

构建首个大规模口语新加坡英语语料库,提升语音理解能力

Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models

  • 构建MNSC语料库,覆盖语音识别、问答等多任务
  • 提出SingAudioLLM模型,在多项任务上领先10%-30%
  • 为研究新加坡英语方言提供可复用的标准化数据与工具

新加坡英语(Singlish)是一种基于英语的克里奥尔语,在多语言、多文化语境中具有重要研究价值。然而其口语形式长期缺乏系统研究,限制了对其语言结构及应用的理解。为此,我们对最大规模的口语新加坡英语语料库进行标准化与标注,推出多任务国家语音语料库(Multitask National Speech Corpus, MNSC),支持自动语音识别(ASR)、口语问答(SQA)、口语对话摘要(SDS)和副语言问答(PQA)等任务。我们发布标准化数据划分和人工验证的测试集,推动后续研究。同时,提出SingAudioLLM——一种多任务多模态大模型,利用多模态大语言模型并行处理上述任务。实验表明,该模型在新加坡英语语境下具备良好适应性,性能达到当前最优,相比其他AudioLLMs及级联方案提升10%-30%。

原文摘要 · Abstract (English)

Singlish, a Creole language rooted in English, is a key focus in linguistic research within multilingual and multicultural contexts. However, its spoken form remains underexplored, limiting insights into its linguistic structure and applications. To address this gap, we standardize and annotate the largest spoken Singlish corpus, introducing the Multitask National Speech Corpus (MNSC). These datasets support diverse tasks, including Automatic Speech Recognition (ASR), Spoken Question Answering (SQA), Spoken Dialogue Summarization (SDS), and Paralinguistic Question Answering (PQA). We release standardized splits and a human-verified test set to facilitate further research. Additionally, we propose SingAudioLLM, a multi-task multimodal model leveraging multimodal large language models to handle these tasks concurrently. Experiments reveal our models adaptability to Singlish context, achieving state-of-the-art performance and outperforming prior models by 10-30% in comparison with other AudioLLMs and cascaded solutions.

语音识别多模态新加坡英语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。