呼吁用数据节俭代替盲目扩数据,推动负责任AI发展
Position: Stop Preaching and Start Practising Data Frugality for Responsible Development of AI
- 用子集选择方法减少训练数据量,降低能耗
- 仅轻微损失精度,就可大幅减少碳排放
- 适合关注可持续发展的AI研究者与工程师
本文主张机器学习领域应从倡导数据节俭转向实际践行,以实现负责任的AI发展。长期以来,模型进步被等同于数据规模扩张,虽带来显著性能提升,但边际收益递减,同时能源消耗和碳排放持续攀升。尽管数据节俭理念日益普及,实际应用仍停留于口号。本文通过估算ImageNet-1K下游使用的能耗与碳排放,揭示其环境代价;并实证表明,采用子集选择方法可在几乎不损失准确率的前提下,显著降低训练能耗,并缓解数据集偏差。最后提出可操作建议,推动数据节俭从理念走向实践。
原文摘要 · Abstract (English)
This position paper argues that the machine learning community must move from preaching to practising data frugality for responsible artificial intelligence (AI) development. For too long, progress has been equated with ever-larger datasets, driving remarkable advances but now yielding increasingly diminishing performance gains alongside rising energy use and carbon emissions. While awareness of data frugal approaches has grown, their adoption has remained rhetorical, and data scaling continues to dominate development practice. We argue that this gap between preach and practice must be closed, as continued data scaling entails substantial and under-accounted environmental impacts. To ground our position, we provide indicative estimates of the energy use and carbon emissions associated with the downstream use of ImageNet-1K. We then present empirical evidence that data frugality is both practical and beneficial, demonstrating that subset selection methods can substantially reduce training energy consumption with little loss in accuracy, while also mitigating dataset bias. Finally, we outline actionable recommendations for moving data frugality from rhetorical preaching to concrete practice for responsible development of AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。