巴西语言多样性正受生成式AI威胁,因主流模型依赖有文档记录的语言。
Diversidade linguística e inclusão digital: desafios para uma ia brasileira
- 从社会语言学视角分析技术应用导致的语言选择偏差
- 主流大模型训练依赖有文档的语言,形成语言霸权循环
- 关注南美本土语言保护,适合关心AI公平性的研究者
语言多样性是人类的固有属性,随着生成式人工智能的发展正面临威胁。本文基于社会语言学的贡献,探讨了技术应用中语言选择偏差带来的后果,以及一种恶性循环:某种语言因拥有足够语言文档而被标准化并占据主导地位,进而进一步强化其在大语言模型训练中的优先性,导致其他语言被边缘化。
原文摘要 · Abstract (English)
Linguistic diversity is a human attribute which, with the advance of generative AIs, is coming under threat. This paper, based on the contributions of sociolinguistics, examines the consequences of the variety selection bias imposed by technological applications and the vicious circle of preserving a variety that becomes dominant and standardized because it has linguistic documentation to feed the large language models for machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。