arXiv:2609.06963cs.CL2026-09

构建土耳其语-英语混用语料库,推动多语言识别研究

TurEngMix: A Text Corpus and Benchmark for Turkish-English Code-Mixed Language Identification and Named Entity Recognition

论文配图:TurEngMix: A Text Corpus and Benchmark for Turkish-English Code-Mixed Language Identification and Named Entity Recognition
图 1 · 摘自论文原文
  • 收集5.5千条社交媒体混用文本,构建真实场景语料库
  • 混合词标注错误率高达单语的5.2至6.3倍
  • 适合关注低资源语言、混合语处理的研究者

自然语言处理系统在混用语言文本上表现不佳,尤其对低资源语言对。土耳其语-英语混用尤为特殊:英语词干与土耳其语后缀结合形成单一混合词。我们提出TurEngMix,一个包含5.5K条噪声自然社交文本(486,974个词元)的语料库,富含土耳其语-英语混用特征。基于此,我们构建新的土耳其语-英语混用语言识别(LID)与命名实体识别(NER)基准,包含15,000个专家标注词元。评估解码器大模型与微调编码器基线发现,单语词元识别可靠,但所有模型在混合词元上的错误率极高。对于形态融合词元,GPT-4o与Qwen的NER错误率分别高出5.2倍和6.3倍。这凸显形态整合仍是核心挑战。我们已开源语料库、标注数据与代码,支持未来计算与社会语言学研究。

原文摘要 · Abstract (English)

Natural language processing systems underperform on code-mixed text, particularly for low-resource language pairs. Turkish-English poses a further challenge: it lets English stems combine with Turkish suffixes to form single mixed-language tokens. We introduce TurEngMix, a corpus of 5.5K noisy, naturally occurring social media posts (486,974 tokens) rich in Turkish-English code-mixing. From this corpus, we construct a new Turkish-English benchmark for code-mixed language identification (LID) and named entity recognition (NER), comprising 15K expert-annotated tokens. Evaluating both decoder LLM and fine-tuned encoder baselines, we find that monolingual Turkish and English tokens are labeled reliably, but all models have high error rates on mixed-language tokens for both LID and NER. For morphologically integrated tokens, NER error rates were 5.2x and 6.3x higher for GPT-4o and Qwen, respectively. This highlights how morphological integration remains a challenge. We release the corpus, annotations, and code to support future computational and sociolinguistic research on Turkish-English code-mixing.

代码混用语言识别命名实体识别多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。