arXiv:2510.26254cs.CL2025-10中稿 · LREC 2026被引 2

大模型识别借词能力差,可能加剧少数语言消亡

Are Language Models Borrowing-Blind? A Multilingual Evaluation of Loanword Identification across 10 Languages

  • 在10种语言上测试大模型借词识别能力
  • 模型区分借词与本族词准确率普遍偏低
  • 研究提醒需关注少数语言的保护与工具开发

语言历史中,词汇常从一种语言借入另一种语言并逐渐融入。母语者通常能区分借词与原生词汇,尤其在双语社区中,强势语言持续向弱势语言输入词汇。本文评估了多种预训练语言模型在10种语言上的借词识别能力。尽管提供明确指令和上下文信息,模型仍难以有效区分借词与本地词汇。结果印证了现有研究:现代NLP系统存在对借词的偏好倾向。该工作对开发少数语言的NLP工具、支持受强势语言影响的社区语言保护具有重要意义。

原文摘要 · Abstract (English)

Throughout language history, words are borrowed from one language to another and gradually become integrated into the recipient's lexicon. Speakers can often differentiate these loanwords from native vocabulary, particularly in bilingual communities where a dominant language continuously imposes lexical items on a minority language. This paper investigates whether pretrained language models, including large language models, possess similar capabilities for loanword identification. We evaluate multiple models across 10 languages. Despite explicit instructions and contextual information, our results show that models perform poorly in distinguishing loanwords from native ones. These findings corroborate previous evidence that modern NLP systems exhibit a bias toward loanwords rather than native equivalents. Our work has implications for developing NLP tools for minority languages and supporting language preservation in communities under lexical pressure from dominant languages.

语言模型借词识别少数语言NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。