四大语言模型对书写系统支持不均,根源在于帝国历史遗留的数字不平等。
The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

- 构建七维数字书写支持指数,量化300种文字的数字化程度。
- 仅9.7%文字获完全支持,45种文字分词效率相差31.7倍,误差高度趋同。
- 模型共性错误源于历史帝国影响,非单一模型设计导致,适合关注公平性的研究者。
大型语言模型在处理全球书写系统时存在严重不平等。我们构建了数字书写表征指数(DSRI),从七个维度衡量数字支持水平,并应用于全球书写数据库(Fukui, 2026)中的300种书写系统。仅有29种(9.7%)获得现代数字基础设施的完整支持;在158种现存文字中,60种(38.0%)缺乏完整支持。45种文字的并行文本测试显示,分词效率差异达31.7倍。序列中介模型表明:帝国干预→使用者人口→网络语料→分词效率的路径具有完全中介效应,帝国直接影响不显著(beta = -0.22, p = 0.39),结构方程模型拟合指标接近饱和(n = 45);校正偏倚的自助法置信区间逼近零,故将中介关系视为提示性而非确认性。在四个独立大模型家族(Claude、GPT-4o、Grok、DeepSeek;共12,000次API调用)中,基线偏差误差模式呈现高度收敛(斯皮尔曼相关系数rho = 0.85–0.98,全部p < 0.002)。172项文字特征被所有四模型一致答错;过度归因是低估的3.9倍,其中“用于宗教”一项集中了43.6%的共性错误(富集度4.1倍)。剔除宗教因素后,跨架构收敛仍保持(九项特征平均rho = 0.87),过度归因不对称性依旧显著(1.77:1,二项检验p = 0.008),表明偏差为多通道而非单通道所致。结果支持一种解释:历史上帝国对书写群体造成的结构性不平等,通过共享训练语料延续至当代语言模型,而非由个别模型设计造成。
原文摘要 · Abstract (English)
Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a seven-axis measure of digital support, and applied it to the 300 writing systems of the Global Script Database (Fukui, 2026). Only 29 scripts (9.7%) are fully supported by contemporary digital infrastructure; among 158 living scripts, 60 (38.0%) lack complete support. Tokenizer efficiency varies by a factor of 31.7 across 45 scripts measured with parallel text. A serial mediation model -- imperial intervention to speaker population to web corpus to tokenizer efficiency -- is consistent with full mediation, with the direct effect of empire indistinguishable from zero (beta = -0.22, p = 0.39) and structural equation model fit indices indistinguishable from saturation at n = 45; the bias-corrected bootstrap CI grazes zero, and we treat the mediation as suggestive rather than confirmatory. Across four independent LLM families (Claude, GPT-4o, Grok, DeepSeek; 12,000 API calls), base-rate-deviation error patterns converge at Spearman rho = 0.85-0.98 (all p < 0.002). 172 script-feature items are answered identically wrong by all four models; over-attribution outnumbers under-recognition 3.9:1, and "used for religion" alone concentrates 43.6% of convergent errors (enrichment 4.1x). With religion excluded as a sensitivity check, the cross-architecture convergence is preserved (mean rho = 0.87 on nine features) and the over-attribution asymmetry persists at 1.77:1 (n = 97, binomial p = 0.008), indicating multi-channeled rather than single-channeled bias. The findings are consistent with an interpretation in which the structural inequalities historical empires inflicted on script communities persist in contemporary language models through the shared training corpus rather than through any individual model's design choices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。