arXiv:2511.08637cs.CYcs.AI2025-11AAAI被引 5

研究网页数据训练AI时,主人如何表达同意,发现大量数据有版权提示但未被尊重。

How Do Data Owners Say No? A Case Study of Data Consent Mechanisms in Web-Scraped Vision-Language AI Training Datasets

  • 分析数据集中的版权声明、水印和网站协议,识别数据主人的同意信号。
  • 12.8亿样本中至少1.22亿带版权提示,前50大域名60%内容禁止抓取。
  • 9-13%样本含水印,现有检测方法无法准确识别,凸显合规漏洞。

互联网已成为现代文生图及视觉语言模型训练的主要数据来源,但大规模数据采集是否尊重数据所有者意愿日益模糊。忽视数据使用同意不仅引发伦理争议,更已升级为版权侵权诉讼。本文以包含128亿文本-图像对的DataComp数据集为例,研究其中数据主人的同意表达方式。我们从样本级(如版权声明、水印、元数据)与网站域级(如服务条款ToS、Robots协议)两方面分析,发现CommonPool中至少1.22亿样本存在版权提示,前50个域名中60%的内容来自禁止抓取的网站。同时估计9-13%(置信区间)的样本含水印,但现有检测方法无法高保真识别。结果表明,数据主人通过多种渠道传达同意,而当前AI数据收集流程未能充分尊重。这揭示了现有数据集构建与发布实践的局限性,亟需建立考虑AI用途的统一数据同意框架。

原文摘要 · Abstract (English)

The internet has become the main source of data to train modern text-to-image or vision-language models, yet it is increasingly unclear whether web-scale data collection practices for training AI systems adequately respect data owners' wishes. Ignoring the owner's indication of consent around data usage not only raises ethical concerns but also has recently been elevated into lawsuits around copyright infringement cases. In this work, we aim to reveal information about data owners' consent to AI scraping and training, and study how it's expressed in DataComp, a popular dataset of 12.8 billion text-image pairs. We examine both the sample-level information, including the copyright notice, watermarking, and metadata, and the web-domain-level information, such as a site's Terms of Service (ToS) and Robots Exclusion Protocol. We estimate at least 122M of samples exhibit some indication of copyright notice in CommonPool, and find that 60\% of the samples in the top 50 domains come from websites with ToS that prohibit scraping. Furthermore, we estimate 9-13\% with 95\% confidence interval of samples from CommonPool to contain watermarks, where existing watermark detection methods fail to capture them in high fidelity. Our holistic methods and findings show that data owners rely on various channels to convey data consent, of which current AI data collection pipelines do not entirely respect. These findings highlight the limitations of the current dataset curation/release practice and the need for a unified data consent framework taking AI purposes into consideration.

数据合规版权问题AI训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。