区分借词与代码切换,关键在标注边界而非模型本身。
Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification
- 标注指南明确保留借词为单一语言,仅将句级切换标为混合。
- 不同模型表现差异大,但根本瓶颈是标注标准不统一。
- 适合关注多语言文本标注规范的研究者与实践者。
现成的语音识别(LID)工具和字符启发法会错误地将俄语借词标记为混合语言:在共用西里尔字母背景下,俄语词汇嵌入哈萨克语文本看似代码切换。我们发布了文档级别的黄金标注LID数据集,其标注指南将融合借词视为哈萨克语,仅将句级切换标记为混合,并设计了在LID后使用混合文本的情感分析池,采用先过滤再识别的级联策略。在共享测试集上,FastText、Lingua、原始与窗口化HeLI、字符三元组NB及XLM-R模型的表现从弱到强不等。性能差距表明瓶颈在于借词与代码切换的标注边界,而非模型类别本身。
原文摘要 · Abstract (English)
Off-the-shelf LID and letter heuristics over-label Kazakh-Russian social text as mixed: Russian loanwords inside Kazakh look like code-switching under a shared Cyrillic script. We release a document-level gold LID set whose guideline keeps integrated borrowings as Kazakh and reserves mixed for clause-level switches, plus a mixed-only sentiment pool used after LID in a filter-first cascade. On a shared LID test, FastText, Lingua, raw and windowed HeLI, character-trigram NB, and XLM-R range from weak to strong performance. The gap shows the bottleneck is the loanword-vs-switch annotation boundary, not model class alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。