arXiv:2603.10789cs.CL2026-03中稿 · LREC2026, 4 Figure…被引 1

分析卢森堡语27年新闻中的借词现象,揭示语言混合的深层模式。

LuxBorrow: From Pompier to Pompjee, Tracing Borrowing in Luxembourgish

  • 构建分层识别与解析流程,精准定位卢森堡语中的借词来源。
  • 发现77.1%文章含至少一种外来语,但混合程度整体较浅,中位数仅7.00。
  • 法语主导借词适应,且借词形态调整趋势随时间增强,适合语言演化研究者。

我们提出LuxBorrow,对1999至2025年间卢森堡语新闻进行首次以借词为核心的分析,涵盖259,305篇RTL文章和4370万词元。该方法结合句子级语言识别(卢/德/法/英)与词元级借词解析器,仅作用于卢森堡语句子,利用词干化、借词词表及编译的形态与拼写规则。实证显示,卢森堡语始终为基底语言,多语实践普遍:77.1%的文章包含至少一种源语言,65.4%使用三种或四种语言。广度不等于强度:中位数代码混杂指数(CMI)从3.90(卢+1)升至7.00(卢+3),表明借词集中于局部而非均衡双语文本。领域与时期分析显示,CMI从1999-2007年的6.1增至2020年峰值8.4。共发现25,444次词元级适应,其中形态调整占63.8%,拼写调整占35.9%,词汇调整仅0.3%。最常见拼写规则如on->oun、eur->er;形态规则整体主导。历时分析显示,代码切换加剧,形态适配借词从少量增长。法语是主要来源,德语小幅增长,英语影响微弱。建议采用借词中心评估体系,包括借用词元率、借用类型率、源语言熵及同化比率,而非仅依赖文档级混杂指数。

原文摘要 · Abstract (English)

We present LuxBorrow, a borrowing-first analysis of Luxembourgish (LU) news spanning 27 years (1999-2025), covering 259,305 RTL articles and 43.7M tokens. Our pipeline combines sentence-level language identification (LU/DE/FR/EN) with a token-level borrowing resolver restricted to LU sentences, using lemmatization, a collected loanword registry, and compiled morphological and orthographic rules. Empirically, LU remains the matrix language across all documents, while multilingual practice is pervasive: 77.1% of articles include at least one donor language and 65.4% use three or four. Breadth does not imply intensity: median code-mixing index (CMI) increases from 3.90 (LU+1) to only 7.00 (LU+3), indicating localized insertions rather than balanced bilingual text. Domain and period summaries show moderate but persistent mixing, with CMI rising from 6.1 (1999-2007) to a peak of 8.4 in 2020. Token-level adaptations total 25,444 instances and exhibit a mixed profile: morphological 63.8%, orthographic 35.9%, lexical 0.3%. The most frequent individual rules are orthographic, such as on->oun and eur->er, while morphology is collectively dominant. Diachronically, code-switching intensifies, and morphologically adapted borrowings grow from a small base. French overwhelmingly supplies adapted items, with modest growth for German and negligible English. We advocate borrowing-centric evaluation, including borrowed token and type rates, donor entropy over borrowed items, and assimilation ratios, rather than relying only on document-level mixing indices.

语言演化借词分析多语混合自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。