探索用户自愿贡献数据的生成式AI如何实现更公平的模型发展
Can Generative AI be Egalitarian?
- 借鉴维基百科模式,用自愿协作内容训练AI模型
- 相比掠夺性数据采集,该模式提升数据多样性与社会契合度
- 适合关注AI伦理、可持续发展的研究者与实践者
生成式AI的迅猛发展依赖于对网络内容的大规模价值提取,往往缺乏相应回馈,这延续并加剧了监视资本主义的剥削模式。同时,巨大的利润潜力也削弱了科技公司对负责任AI的承诺,引发重大伦理与社会问题。本文探讨一种替代路径:基于用户自愿、协作提供的内容构建模型。受维基百科成功经验启发,该“平等主义”方法可能使未来基础模型在设计、开发与约束上更具社会包容性。我们主张此模式不仅伦理上更合理,还能使模型更贴合用户需求、训练数据更丰富,并更好地契合社会价值观。同时,文章分析了该模式面临的挑战,包括可扩展性、质量控制以及志愿者内容固有的偏见问题。
原文摘要 · Abstract (English)
The recent explosion of "foundation" generative AI models has been built upon the extensive extraction of value from online sources, often without corresponding reciprocation. This pattern mirrors and intensifies the extractive practices of surveillance capitalism, while the potential for enormous profit has challenged technology organizations' commitments to responsible AI practices, raising significant ethical and societal concerns. However, a promising alternative is emerging: the development of models that rely on content willingly and collaboratively provided by users. This article explores this "egalitarian" approach to generative AI, taking inspiration from the successful model of Wikipedia. We explore the potential implications of this approach for the design, development, and constraints of future foundation models. We argue that such an approach is not only ethically sound but may also lead to models that are more responsive to user needs, more diverse in their training data, and ultimately more aligned with societal values. Furthermore, we explore potential challenges and limitations of this approach, including issues of scalability, quality control, and potential biases inherent in volunteer-contributed content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。