arXiv:2608.00581cs.CLcs.AI2026-08

区分借词与代码切换,关键在标注边界而非模型本身。

Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification

  • 标注指南明确保留借词为单一语言,仅将句级切换标为混合。
  • 不同模型表现差异大,但根本瓶颈是标注标准不统一。
  • 适合关注多语言文本标注规范的研究者与实践者。

现成的语音识别(LID)工具和字符启发法会错误地将俄语借词标记为混合语言:在共用西里尔字母背景下,俄语词汇嵌入哈萨克语文本看似代码切换。我们发布了文档级别的黄金标注LID数据集,其标注指南将融合借词视为哈萨克语,仅将句级切换标记为混合,并设计了在LID后使用混合文本的情感分析池,采用先过滤再识别的级联策略。在共享测试集上,FastText、Lingua、原始与窗口化HeLI、字符三元组NB及XLM-R模型的表现从弱到强不等。性能差距表明瓶颈在于借词与代码切换的标注边界,而非模型类别本身。

原文摘要 · Abstract (English)

Off-the-shelf LID and letter heuristics over-label Kazakh-Russian social text as mixed: Russian loanwords inside Kazakh look like code-switching under a shared Cyrillic script. We release a document-level gold LID set whose guideline keeps integrated borrowings as Kazakh and reserves mixed for clause-level switches, plus a mixed-only sentiment pool used after LID in a filter-first cascade. On a shared LID test, FastText, Lingua, raw and windowed HeLI, character-trigram NB, and XLM-R range from weak to strong performance. The gap shows the bottleneck is the loanword-vs-switch annotation boundary, not model class alone.

代码切换语言识别标注标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。