揭示AI基础设施对小语种使用者的系统性排斥,提出需以公平为导向重构技术设计。
Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

- 分析孟加拉语在网页内容、训练数据、分词机制和网络接入四方面的结构性缺陷
- 孟加拉语仅占全球网页内容0.5%,但人口占比近4%;英孟训练数据量差距达67:1
- 倡导离线优先设计,将语言公平纳入AI基础设施核心考量
面向教育与语言支持的人工智能工具常被视作解决资源匮乏社区访问差距的可扩展方案。然而,这些工具背后的基础设施——包括训练语料、分词方案、评估基准和部署架构——在模型训练前便系统性地排斥小语种使用者。本文以全球使用人数最多的语言之一孟加拉语为例,聚焦低网络连接环境下的AI辅助教育。研究发现四大相互关联的失败:网页存在度严重不足(孟加拉语占全球网页内容不到0.5%,但人口占比近4%);主流多语语料中英语与孟加拉语的训练标记数量差距达67:1;其音节字母文字特性导致分词时标记密度更高,加剧数据短缺;农村地区互联网渗透率仅36.5%,远低于城市71.4%。这些缺陷反映了长期资源分配决策、机构优先级与设计默认值未以小语种为中心。我们主张将数据稀缺视为结构性障碍而非孤立技术问题,并将离线优先设计作为公平导向的基础设施策略。最后提出语言学与AI研究应致力于减少此类结构性不平等。
原文摘要 · Abstract (English)
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds the data deficit through higher token fertility; and connectivity exclusion, with individual internet penetration at 36.5% in rural areas compared with 71.4% in urban areas. These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development. We argue that dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation, and that offline-first design should be treated as an equity-oriented infrastructure strategy. We conclude with directions for linguistics and AI research aimed at reducing these structural inequalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。