arXiv:2509.24310eess.AS2025-09中稿 · IEEE TASLP被引 7

提出新方法生成更自然的代码转换语音,提升跨语言语音识别效果

Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives

  • 用简化等价约束理论引导大模型生成真实感代码转换文本
  • 结合语音合成增强数据,使识别准确率显著提升
  • 强调需根据语言特征匹配模型与数据策略,适配多语种场景

代码转换自动语音识别(CS-ASR)因句内语言混杂和口音干扰导致语音边界模糊而面临挑战。尽管单语种资源丰富,但标注的代码转换数据仍稀缺。本文从模型与数据双视角系统分析该问题:对比先进算法如语言特异性处理与多任务学习的有效性;探索语音合成(TTS)作为数据增强手段,研究语言混淆与口音偏见的影响;进一步提出简化等价约束理论(SECT),指导大语言模型生成符合语言规律的代码转换文本。SECT在语音识别性能与语言质量评估中均优于现有方法,生成内容更贴近真实语境。利用SECT生成语音-文本对进行TTS合成后,显著提升了CS-ASR表现。研究表明,有效实现CS-ASR需依据具体语言特征协调模型与数据策略。

原文摘要 · Abstract (English)

Code-switching automatic speech recognition (CS-ASR) presents unique challenges due to language confusion introduced by spontaneous intra-sentence switching and accent bias that blurs the phonetic boundaries. Although the constituent languages may be individually high-resource, the scarcity of annotated code-switching data further compounds these challenges. In this paper, we systematically analyze CS-ASR from both model-centric and data-centric perspectives. By comparing state-of-the-art algorithmic methods, including language-specific processing and auxiliary language-aware multi-task learning, we discuss their varying effectiveness across datasets with different linguistic characteristics. On the data side, we first investigate TTS as a data augmentation method. By varying the textual characteristics and speaker accents, we analyze the impact of language confusion and accent bias on CS-ASR. To further mitigate data scarcity and enhance textual diversity, we propose a prompting strategy by simplifying the equivalence constraint theory (SECT) to guide large language models (LLMs) in generating linguistically valid code-switching text. The proposed SECT outperforms existing methods in ASR performance and linguistic quality assessments, generating code-switching text that more closely resembles real-world code-switching text. When used to generate speech-text pairs via TTS, SECT proves effective in improving CS-ASR performance. Our analysis of both model- and data-centric methods underscores that effective CS-ASR requires strategies to be carefully aligned with the specific linguistic characteristics of the code-switching data.

语音识别代码转换大模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。