AI基准测试实则固化权力结构,加剧科研不公。
The Benchmark Trap: Structures of Power and Injustice in AI Evaluations
- 用社会理论分析基准测试如何塑造研究格局
- 揭示其正强化现有巨头的资源垄断
- 适合关注AI伦理与公平性的研究者阅读
人工智能基准测试并非中立评估工具,而是塑造竞争、权力与研究方向的社技术物。它们通过标准化评估和排行榜机制,将声望、引用、信任与机构影响力奖励给顶尖性能系统。随着开发先进AI系统的成本上升,这些奖励日益集中于有产业资金支持的大型实验室。本文基于伊里斯·马里昂·杨的压迫与结构性不公理论,指出当前基准测试实践可能系统性地伤害多元参与者,契合其提出的四种'压迫面孔'。基准测试文化被视为结构性不公的来源,因其危害源于被普遍接受、个体可辩护的常规做法与网络效应,即便无明确恶意亦然。它强化既有权力结构,限制研究路径,反而阻碍了认知上稳健且对社会有益的进步。
原文摘要 · Abstract (English)
Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmarks standardise the assessment of systems and facilitate the creation of leaderboards that reward state-of-the-art performance with prestige, citations, trust, and institutional influence. As the costs of developing competitive AI systems rise, these rewards increasingly concentrate among powerful, industry-funded labs. This paper situates these concerns within Iris Marion Young's theories of oppression and structural injustice. It argues that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Young's "faces of oppression". Benchmarking culture is further framed as a source of structural injustice, as these harms emerge from normalised, individually defensible practices and network effects, even without explicit wrongdoing. By reinforcing existing power structures and narrowing possible research trajectories, benchmarking may in fact prevent the field from advancing in epistemically robust and socially beneficial ways.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。