arXiv:2603.28213cs.CL2026-03

探究大模型对非标准语言的偏见,推动技术公平与去殖民化

\textit{Versteasch du mi?} Computational and Socio-Linguistic Perspectives on GenAI, LLMs, and Non-Standard Language

  • 从技术与社会语言学双重视角分析大模型对非标准语言的歧视
  • 以南蒂罗尔方言和库尔德语为例,揭示语言多样性在算法中的缺失
  • 呼吁构建更包容的数字语言策略,适合政策制定者与技术伦理研究者

大型语言模型(LLMs)与生成式人工智能(GenAI)被证实对低使用率语言存在不公平现象,并加剧数字语言鸿沟。已有批判性社会语言学研究指出,这些技术依赖于历史上以欧洲民族主义和殖民项目为基础的语言标准化进程,同时强化了语言为“单一、单语、语法标准化系统”的认知。本文结合技术与语言政策交叉研究,基于作者在批判社会语言学与计算语言学的专业背景,选取南蒂罗尔方言(广泛用于意大利南蒂罗尔地区的非正式交流)及库尔德语变体作为案例,开展跨学科探讨。文章讨论从技术角度如何让大模型处理非标准语言,并反思其是否能促成“民主与去殖民化的数字及机器学习策略”,具有直接政策意义。

原文摘要 · Abstract (English)

The design of Large Language Models (LLMs) and generative artificial intelligence (GenAI) has been shown to be "unfair" to less-spoken languages (Petrov et al., 2023) and to deepen the digital language divide (Bella et al., 2023). Critical sociolinguistic work has also argued that these technologies are not only made possible by prior sociohistorical processes of linguistic standardisation, often grounded in European nationalist and colonial projects (Migge and Schneider, 2025), but also exacerbate epistemologies of language as "monolithic, monolingual, syntactically standardized systems of meaning" (Schneider, 2024, p. 5). In our paper, we draw on earlier work on the intersections of technology and language policy (Kelly-Holmes, 2019) and bring our respective expertise in critical sociolinguistics and computational linguistics to bear on an interrogation of these arguments. We take two different complexes of non-standard linguistic varieties in our respective repertoires-South Tyrolean dialects, which are widely used in informal communication in South Tyrol, Italy (Alber et al., 2024), as well as varieties of Kurdish-as starting points to an interdisciplinary exploration of the intersections between GenAI and linguistic variation and standardisation. We discuss both how LLMs can be made to deal with non-standard language from a technical perspective, and whether, when or how this can contribute to "democratic and decolonial digital and machine learning strategies" (Migge and Schneider, 2025, p. 12), which has direct policy implications.

语言公平大模型去殖民化非标准语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。