arXiv:2508.15244cs.CLcs.SD2025-08EMNLP被引 2

用替换同义词方法生成自然语码转换语音,解决多语言语音技术数据不足问题。

UniCoM: A Universal Code-Switching Speech Generator

  • 通过同义词替换算法生成语码转换语音,保持原意不变。
  • 构建的CS-FLEURS数据集在可懂度和自然度上表现优异。
  • 适合做语音识别与跨语言语音翻译的研究者使用。

语码转换(CS)指单个说话者在话语中交替使用两种或以上语言,是真实对话中的常见现象,对多语言语音技术构成重大挑战。然而,能够处理该现象的系统仍研究不足,主要受限于合适数据集的缺乏。为此,我们提出通用语码混合器(UniCoM),一种无需改变句子语义即可生成高质量、自然语码转换语音的新流程。该方法采用名为SWORDS(Substituting WORDs with Synonyms)的算法,通过考虑词性将选定词汇替换为对应译文来生成语码转换语音。利用UniCoM,我们构建了语码转换FLEURS(CS-FLEURS)数据集,专用于自动语音识别(ASR)与语音到文本翻译(S2TT)。实验表明,CS-FLEURS在客观与主观评估指标上均表现良好,其可懂度与自然度与现有数据集相当。我们期望该方法能推动语码转换语音技术发展,助力更包容的多语言系统建设。

原文摘要 · Abstract (English)

Code-switching (CS), the alternation between two or more languages within a single speaker's utterances, is common in real-world conversations and poses significant challenges for multilingual speech technology. However, systems capable of handling this phenomenon remain underexplored, primarily due to the scarcity of suitable datasets. To resolve this issue, we propose Universal Code-Mixer (UniCoM), a novel pipeline for generating high-quality, natural CS samples without altering sentence semantics. Our approach utilizes an algorithm we call Substituting WORDs with Synonyms (SWORDS), which generates CS speech by replacing selected words with their translations while considering their parts of speech. Using UniCoM, we construct Code-Switching FLEURS (CS-FLEURS), a multilingual CS corpus designed for automatic speech recognition (ASR) and speech-to-text translation (S2TT). Experimental results show that CS-FLEURS achieves high intelligibility and naturalness, performing comparably to existing datasets on both objective and subjective metrics. We expect our approach to advance CS speech technology and enable more inclusive multilingual systems.

语码转换语音生成多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。