研究网页爬虫退出对大模型性能的影响,发现通用模型不受影响,专业领域则需高质量版权数据。
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
- 提出数据合规差距(DCG)概念,量化合规数据与非合规数据训练模型的性能差异。
- 1.5B模型实验显示,通用知识获取几乎无损失(接近0% DCG),但生物医学等专业领域性能下降。
- 适合关注AI伦理、数据合规与模型性能平衡的研究者和政策制定者参考。
随着内容版权方越来越多地采用网页爬虫退出机制,大语言模型(LLM)的数据合规性对模型性能的影响成为关键问题。本文提出“数据合规差距”(DCG)概念,用于量化在遵守网页爬虫退出规则的语料库上训练的模型与未遵守规则的模型之间的性能差异。我们在两种场景下进行评估:从头训练模型,以及对已有合规模型进行持续预训练(模拟后期加入受版权保护数据的情形)。使用1.5B参数模型的实验表明,截至2025年1月,遵守网络数据退出规则不会显著影响通用知识的获取(接近0% DCG)。然而,在生物医学研究等专业领域,排除主要出版商的数据会导致性能下降。结果表明,通用型大模型可仅用开放数据训练并保持同等性能,但在特定领域,后期引入高质量版权数据仍有益处。本研究为长期争论的数据合规与模型性能之间的权衡提供了实证依据,有助于未来人工智能训练实践与政策讨论。
原文摘要 · Abstract (English)
The increasing adoption of web crawling opt-outs by copyright holders of online content raises critical questions about the impact of data compliance on large language model (LLM) performance. However, little is known about how these restrictions (and the resultant filtering of pretraining datasets) affect the capabilities of models trained using these corpora. In this work, we conceptualize this effect as the $\textit{data compliance gap}$ (DCG), which quantifies the performance difference between models trained on datasets that comply with web crawling opt-outs, and those that do not. We measure the data compliance gap in two settings: pretraining models from scratch and continual pretraining from existing compliant models (simulating a setting where copyrighted data could be integrated later in pretraining). Our experiments with 1.5B models show that, as of January 2025, compliance with web data opt-outs does not degrade general knowledge acquisition (close to 0\% DCG). However, in specialized domains such as biomedical research, excluding major publishers leads to performance declines. These findings suggest that while general-purpose LLMs can be trained to perform equally well using fully open data, performance in specialized domains may benefit from access to high-quality copyrighted sources later in training. Our study provides empirical insights into the long-debated trade-off between data compliance and downstream model performance, informing future discussions on AI training practices and policy decisions. Our website is available at https://data-compliance.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。