警惕生成式AI带来的数字垃圾,推动可持续发展
Responsible Data Stewardship: Generative AI and the Digital Waste Problem
- 将未使用数据定义为数字垃圾,提出可持续发展的伦理要求
- 分析跨学科管理经验,提出技术与文化双轨减废策略
- 适合关注AI环保、长期可持续性的研究者和开发者
随着生成式AI广泛应用,文本、图像、音频和视频等模态的合成数据产量空前增长。尽管模型训练与推理的能耗已受关注,但一个关键的可持续性挑战仍被忽视:数字垃圾——指存储但无明确(或即时)用途的数据所消耗的资源。本文首次在人工智能语境中引入该术语,将数字垃圾视为生成式AI开发中的伦理责任,强调环境可持续性是负责任创新的核心。借鉴其他领域的数字资源管理实践,我们识别出可迁移的方法,并提出涵盖研究方向、技术干预与文化转变的具体建议,以缓解无限数据存储带来的环境影响。通过将AI伦理从偏见与隐私扩展至代际环境正义,本工作构建了更全面的伦理框架,涵盖生成式AI全生命周期的环境影响。
原文摘要 · Abstract (English)
As generative AI systems become widely adopted, they enable unprecedented creation levels of synthetic data across text, images, audio, and video modalities. While research has addressed the energy consumption of model training and inference, a critical sustainability challenge remains understudied: digital waste. This term refers to stored data that consumes resources without serving a specific (and/or immediate) purpose. This paper presents this terminology in the AI context and introduces digital waste as an ethical imperative within (generative) AI development, positioning environmental sustainability as core for responsible innovation. Drawing from established digital resource management approaches, we examine how other disciplines manage digital waste and identify transferable approaches for the AI community. We propose specific recommendations encompassing re-search directions, technical interventions, and cultural shifts to mitigate the environmental consequences of in-definite data storage. By expanding AI ethics beyond immediate concerns like bias and privacy to include inter-generational environmental justice, this work contributes to a more comprehensive ethical framework that considers the complete lifecycle impact of generative AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。