模型坍缩正威胁低资源社区的AI民主化,需警惕其数据污染与不公。
Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities

- 用前代模型输出训练新模型,导致性能下降与数据失真。
- 模型坍缩使长尾数据更难获取,加剧边缘群体被忽视。
- 呼吁关注低资源群体,推动更公平的AI发展路径。
模型坍缩指生成模型在基于先前模型输出的数据上训练时性能退化的问题,随着人工生成内容泛滥而日益严峻。大型语言模型常复现训练数据中的高频模式,依赖海量数据,且环境成本高昂。这些因素共同导致数据质量下降、文化偏见强化及资源浪费。本文结合多方观点,指出模型坍缩正威胁人工智能的普惠化进程:它降低训练效率,扭曲数据分布,尤其损害低资源与边缘化社群。文章探讨其环境与文化影响,定位自身于近期关于模型坍缩的立场论文中,并提出应对方向。最后,建议采取措施缓解其负面影响。
原文摘要 · Abstract (English)
Model collapse, the degradation in performance that arises when generative models are trained on the outputs of prior models, is an increasing concern as artificially generated content proliferates. Related critiques of large language models have highlighted their tendency to reproduce frequent patterns in training data, their reliance on vast datasets, and their substantial environmental cost. Together, these factors contribute to data degradation, the reinforcement of cultural biases, and inefficient resource use. In this position paper we aim to combine these views and argue that model collapse threatens current efforts to democratize AI. By reducing training efficiency and skewing data distributions away from the tails of their support, model collapse disproportionately impacts low-resource and marginalized communities. We examine both the environmental and cultural implications of this phenomenon, situate our position within recent position papers on model collapse, and conclude with a call to action. Finally, we outline initial directions for mitigating these effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。