首个韩英混用语音识别评估框架,助力多语言模型提升跨语种识别能力。
HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
- 构建分层级标注的韩英混用语音数据集,支持词、短语、句级评估
- 多语言模型经合成数据微调后,混用识别准确率显著提升
- 适合关注多语言语音识别与跨语言交互研究的学者使用
尽管多语言自动语音识别(ASR)取得进展,但日常口语中常见的语言混用(代码切换,CS)仍是严重未被充分研究的挑战。本文提出HiKE:首个全球可访问的非合成韩英混用语音识别基准,旨在为多语言ASR模型提供精确评估手段并推动该领域研究。该框架包含高质量、自然的跨话题混用数据,配有精细的借词标注和分层次的代码切换标注体系(词、短语、句子),可系统评估模型在不同层级代码切换下的表现。通过对多种多语言ASR模型的评测及微调实验,结果表明:虽然多数模型初始时在代码切换识别上表现不佳,但通过合成代码切换数据微调后,其性能可显著提升。HiKE已开源:https://github.com/ThetaOne-AI/HiKE。
原文摘要 · Abstract (English)
Despite advances in multilingual automatic speech recognition (ASR), code-switching (CS), the mixing of languages within an utterance common in daily speech, remains a severely underexplored challenge. In this paper, we introduce HiKE: the Hierarchical Korean-English code-switching benchmark, the first globally accessible non-synthetic evaluation framework for Korean-English CS, aiming to provide a means for the precise evaluation of multilingual ASR models and to foster research in the field. The proposed framework not only consists of high-quality, natural CS data across various topics, but also provides meticulous loanword labels and a hierarchical CS-level labeling scheme (word, phrase, and sentence) that together enable a systematic evaluation of a model's ability to handle each distinct level of code-switching. Through evaluations of diverse multilingual ASR models and fine-tuning experiments, this paper demonstrates that although most multilingual ASR models initially exhibit inadequate CS-ASR performance, this capability can be enabled through fine-tuning with synthetic CS data. HiKE is available at https://github.com/ThetaOne-AI/HiKE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。