检测大模型在213国中的偏见,发现南北差距显著
Global AI Bias Audit for Technical Governance
- 用1704个查询测试Llama-3模型在213国的表现
- 仅11.4%回答含可验证数据,且多集中于高收入地区
- 揭示全球南北方技术治理信息鸿沟,适合政策与伦理研究者
本文报告了全球大语言模型(LLM)审计项目的探索性阶段成果。基于全球人工智能数据集(GAID)项目框架,对Llama-3 8B模型进行了跨213个国家、八项技术指标的应力测试,评估其在地理与经济社会层面的技术治理认知偏见。测试涵盖1,704个查询,结果显示模型仅在11.4%的回答中提供可验证的数字/事实信息,且这些信息高度集中于高收入地区。结果表明,低收入国家在技术治理知识上存在系统性信息缺口,加剧了全球南北之间的数字鸿沟。这种差异威胁全球AI安全与包容性治理,使欠发达地区政策制定者可能依赖错误信息或幻觉内容。论文指出当前对齐与训练流程强化了既有地缘经济和地缘政治不平等,呼吁更包容的数据代表,以确保AI真正成为全球公共资源。
原文摘要 · Abstract (English)
This paper presents the outputs of the exploratory phase of a global audit of Large Language Models (LLMs) project. In this exploratory phase, I used the Global AI Dataset (GAID) Project as a framework to stress-test the Llama-3 8B model and evaluate geographic and socioeconomic biases in technical AI governance awareness. By stress-testing the model with 1,704 queries across 213 countries and eight technical metrics, I identified a significant digital barrier and gap separating the Global North and South. The results indicate that the model was only able to provide number/fact responses in 11.4% of its query answers, where the empirical validity of such responses was yet to be verified. The findings reveal that AI's technical knowledge is heavily concentrated in higher-income regions, while lower-income countries from the Global South are subject to disproportionate systemic information gaps. This disparity between the Global North and South poses concerning risks for global AI safety and inclusive governance, as policymakers in underserved regions may lack reliable data-driven insights or be misled by hallucinated facts. This paper concludes that current AI alignment and training processes reinforce existing geoeconomic and geopolitical asymmetries, and urges the need for more inclusive data representation to ensure AI serves as a truly global resource.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。