arXiv:2602.00056cs.CYcs.AI2026-02被引 3

AI发展正从用数据建模转向主动造数据,代价是环境与全球不公加剧。

How Hyper-Datafication Impacts the Sustainability Costs in Frontier AI

  • 提出'超数据化'概念,描述AI主动制造数据的新模式。
  • 分析55万数据集发现数据增长导致能源消耗与碳排放激增。
  • 揭示数据劳动与风险向全球南方转移,呼吁建立责任框架。

过去十年,大规模数据推动了前沿人工智能模型的成功。这一扩张依赖大型科技公司持续汇聚和整理互联网规模的数据集。本文从可持续性视角审视人工智能中大规模数据的环境、社会与经济成本。我们指出,领域正从‘用数据建模’转向‘为建模而创造数据’,称之为‘超数据化’,标志着前沿AI及其社会影响的关键转折点。为量化与语境化数据相关成本,我们分析了来自Hugging Face Hub的大约55万数据集,重点关注数据集增长、存储相关的能源消耗与碳足迹,以及语言数据的社会代表性。我们结合来自肯尼亚数据工作者的定性反馈,考察数据劳动状况,包括大公司直接雇佣及接触暴力内容的风险。此外,通过外部数据源揭示数据中心基础设施的全球不平等。分析表明,超数据化带来了显著且持续上升的环境成本,同时系统性地将劳动风险与表征伤害转移到全球南方。因此,我们提出数据PROOFS建议,涵盖来源追溯、资源意识、所有权、开放性、节俭性与标准,以缓解这些代价。本研究旨在揭示支撑前沿AI的被忽视的数据成本,并激发研究界及更广泛领域的讨论。

原文摘要 · Abstract (English)

Large-scale data has fuelled the success of frontier artificial intelligence (AI) models over the past decade. This expansion has relied on sustained efforts by large technology corporations to aggregate and curate internet-scale datasets. In this work, we examine the environmental, social, and economic costs of large-scale data in AI through a sustainability lens. We argue that the field is shifting from building models from data to actively creating data for building models. We characterise this transition as hyper-datafication, which marks a critical juncture for the future of frontier AI and its societal impacts. To quantify and contextualise data-related costs, we analyse approximately 550,000 datasets from the Hugging Face Hub, focusing on dataset growth, storage-related energy consumption and carbon footprint, and societal representation using language data. We complement this analysis with qualitative responses from data workers in Kenya to examine the labour involved, including direct employment by big tech corporations and exposure to graphic content. We further draw on external data sources to substantiate our findings by illustrating the global disparity in data centre infrastructure. Our analyses reveal that hyper-datafication drives substantial and growing environmental costs while systematically redistributing labour risks and representational harms toward the Global South. Thus, we propose Data PROOFS recommendations spanning provenance, resource awareness, ownership, openness, frugality, and standards to mitigate these costs. Our work aims to make visible the often-overlooked costs of data that underpin frontier AI and to stimulate broader debate within the research community and beyond.

AI伦理数据成本可持续性全球不公

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。