arXiv:2603.24222cs.CL2026-03中稿 · LREC 2026

将语言变异纳入NLP研究,提升模型对真实语言的适应能力

Variation is the Norm: Embracing Sociolinguistics in NLP

  • 构建融合社会语言学的NLP框架,主动引入语言变异
  • 卢森堡语实验显示,变异数据使模型性能下降30%以上
  • 在微调中加入变异数据可显著提升模型鲁棒性,适合多语言研究者

自然语言处理中,语言变异常被视为噪声并被标准化处理,但其实它是语言的核心组成部分。本文提出一个融合社会语言学与NLP技术的框架,主张主动将变异纳入研究设计,反哺模型优化。以正在演变的卢森堡语为例,该语言存在大量拼写变异。实验表明,在包含大量拼写变异的数据上训练和微调的模型,其性能相比标准拼写数据下降超过30%。通过在微调过程中引入变异数据,可有效改善模型表现。该案例强调当前模型对真实语言变异缺乏鲁棒性,而本框架为研究者提供了理论支持与实践路径,推动更贴近现实的语言建模。

原文摘要 · Abstract (English)

In Natural Language Processing (NLP), variation is typically seen as noise and "normalised away" before processing, even though it is an integral part of language. Conversely, studying language variation in social contexts is central to sociolinguistics. We present a framework to combine the sociolinguistic dimension of language with the technical dimension of NLP. We argue that by embracing sociolinguistics, variation can actively be included in a research setup, in turn informing the NLP side. To illustrate this, we provide a case study on Luxembourgish, an evolving language featuring a large amount of orthographic variation, demonstrating how NLP performance is impacted. The results show large discrepancies in the performance of models tested and fine-tuned on data with a large amount of orthographic variation in comparison to data closer to the (orthographic) standard. Furthermore, we provide a possible solution to improve the performance by including variation in the fine-tuning process. This case study highlights the importance of including variation in the research setup, as models are currently not robust to occurring variation. Our framework facilitates the inclusion of variation in the thought-process while also being grounded in the theoretical framework of sociolinguistics.

语言变异社会语言学模型鲁棒性多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。