arXiv:2410.02521cs.CL2024-10EMNLP被引 2

用语法结构识别跨语言说话中的主导语言,比传统语音识别更准。

Methods of Automatic Matrix Language Determination for Code-Switched Speech

  • 基于语法框架分析跨语言语句的主导语言
  • 音频模型在识别主导语言上F1宏得分60%,相关性达0.38
  • 发现非英语语言常作主导语言,与语音识别结果相反

跨语言(CS)是指说话者在不同语言间切换,这一现象在现代日益普遍。为更好描述此类言语,矩阵语言框架(MLF)理论引入了矩阵语言概念,即为跨语言语句提供语法结构的语言。本文基于该理论构建了矩阵语言身份(MLID)判断系统,并将英文/中文、英文/西班牙语跨语言文本与语音的MLID结果与传统的语音语言识别(LID)进行对比。结果显示,基于音频的MLID预测与文本语法规则的相关性更高,且在基于F1宏平均值(60%)和相关性得分(0.38)的识别任务中优于LID。该方法揭示:在跨语言语境中,中文与西班牙语更常作为矩阵语言,而非英语,这与单语语音识别的结果相反。

原文摘要 · Abstract (English)

Code-switching (CS) is the process of speakers interchanging between two or more languages which in the modern world becomes increasingly common. In order to better describe CS speech the Matrix Language Frame (MLF) theory introduces the concept of a Matrix Language, which is the language that provides the grammatical structure for a CS utterance. In this work the MLF theory was used to develop systems for Matrix Language Identity (MLID) determination. The MLID of English/Mandarin and English/Spanish CS text and speech was compared to acoustic language identity (LID), which is a typical way to identify a language in monolingual utterances. MLID predictors from audio show higher correlation with the textual principles than LID in all cases while also outperforming LID in an MLID recognition task based on F1 macro (60%) and correlation score (0.38). This novel approach has identified that non-English languages (Mandarin and Spanish) are preferred over the English language as the ML contrary to the monolingual choice of LID.

跨语言语音识别语言建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。