评测开源大模型在跨国AI治理中的地理偏见,发现低数据国家表现显著更差。
Benchmarking Open-Weight Foundation Models for Global AI Technical Governance
- 用公开权重模型和真实数据集评估各国AI治理知识
- 227国、18指标、近3000个观测点显示地理偏差严重
- 新分类法可区分胡编乱造与诚实拒答,适合政策研究者参考
大语言模型在国家及国际组织的AI治理分析中应用日益广泛,但现有研究表明其对训练数据中代表性不足国家的回答准确性显著下降,即存在地理偏见。现有研究受三方面方法局限:(1) 依赖未公开权重的专有系统,难以独立复现;(2) 评估模型对训练数据截止年后年份的知识,加剧地理信息缺失;(3) 使用粗粒度二分类,无法区分模型自信虚构(HF)与诚实地承认不确定(HR)。本研究通过基准测试四个开源前沿语言模型,使用哈佛数据仓库2026年1月发布的全球AI数据集v2(GAID v2),该数据集包含227个国家的24,453项经验证指标。从中选取18项指标,对应IEEE IRAI 2026框架的八个主题维度,共获得2,990个跨六年的国家-指标-年份观测值(2010–2023)。采用五类分类体系区分:(a) 验证准确(VA)、(b) 自信虚构(HF)、(c) 恳切拒绝(HR)、(d) 定性模糊(QH)、(e) 错误归因(MF)。通过混合效应逻辑回归与差异中的差异(DiD)分析估算地理准确性差异。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in artificial intelligence (AI) governance analysis across national and international organisations. There is, however, growing evidence that such models produce significantly less accurate responses for countries that are underrepresented in their training data-a pattern described in existing literature as geographic bias. Existing studies examining this phenomenon are subject to three methodological limitations that together undermine their findings: (1) reliance on proprietary systems whose weights are not publicly released, which prevents independent replication; (2) evaluation of model knowledge about years that fall after data collection for model training had concluded, leading to geographic ignorance in addition to the natural limits of each model's knowledge; and (3) use of coarse binary response classification that cannot distinguish models' confident fabrication (HF) from their honest acknowledgement of uncertainty. This study addresses all three limitations by benchmarking four open-weight frontier language models against the Global AI Dataset v2 (GAID v2), a verified ground-truth database of 24,453 indicators across 227 countries published on Harvard Dataverse in January 2026. A total of 18 indicators, mapped to the eight thematic dimensions of the IEEE IRAI 2026 framework, are selected from GAID v2, yielding approximately 2,990 country-metric-year observations across six evaluation years (within the period of 2010-2023). Model responses are classified using a five-category scheme that distinguishes (a) verified accuracy (VA), (b) HF, (c) honest refusal (HR), (d) qualitative hedging (QH), and (e) misattribution (MF). Geographic disparities in accuracy are estimated through mixed-effects logistic regression and difference-in-differences (DiD) analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。