arXiv:2608.18117cs.AIecon.GN2026-08

全球AI排行榜忽视南半球,需建立独立治理的区域榜单

Position: AI Leaderboards Are Underserving the Global South: A Case Study from India

  • 以印度为例,提出应建由本地主导的独立治理榜单
  • 现有高质区域基准如IndicSUPERB、MILU未被纳入全球榜单
  • 适合关注AI公平性、区域技术自主的研究者与政策制定者

本文指出,当前人工智能排行榜因缺乏独立治理、利益冲突管理机制及指标演进路径,难以服务全球南方。问题不在于数据缺失——印度已有IndicSUPERB、MILU、LAHAJA等高质量区域基准,非洲有IrokoBench,阿拉伯语有AlGhafa。真正障碍是制度设计:全球榜单未纳入这些基准,也无机制推动其纳入。当北半球付费用户受影响时,商业压力可纠正问题;而南半球缺乏同等影响力。无治理机制下,对印地语、斯瓦希里语或阿拉伯语使用者的影响持续存在且长期被忽视。基于对58位印度AI从业者的调研,普遍支持建立正式治理与披露导向的利益管理机制。解决方案不是更多数据,而是从一开始就构建具备独立治理能力的区域榜单。

原文摘要 · Abstract (English)

This position paper argues that AI leaderboards are structurally ill-suited to serving the Global South because they lack independent governance, conflict-of-interest policies, and mechanisms for metric evolution. The barrier is not missing data; high-quality regional benchmarks already exist: IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic. The barrier is institutional design. Global leaderboards do not include these benchmarks, and no governance mechanism compels them to do so. Commercial pressure corrects leaderboard failures when paying customers in the Global North are affected. The Global South lacks equivalent leverage. Without governance, failures affecting Hindi, Swahili, or Arabic speakers persist indefinitely as documented but unaddressed gaps. Using India as a case study (1.4 billion people, 22 scheduled languages, high-quality benchmarks, but no trusted aggregation), we report findings from a consultation with 58 AI practitioners showing consistent preference for formal governance and disclosure-based conflict management. The solution is not more data but better institutions: regional leaderboards with independent governance from the start.

AI公平性区域基准治理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。