arXiv:2505.00020cs.CLcs.AI2025-05被引 7

测试付费书籍对GPT-4o的识别能力,发现其能识别非公开内容。

Beyond Public Access in LLM Pre-Training Data

  • 用34本付费图书测试模型是否识别非公开内容
  • GPT-4o对非公开内容识别率AUROC达0.82,存在显著识别能力
  • 研究呼吁加强训练数据透明度与内容授权框架

基于34本版权受保护的O'Reilly Media书籍构建合法数据集,采用DE-COP成员推理攻击方法,探究OpenAI大模型是否能识别受版权保护的内容。结果显示,较新且更强的GPT-4o模型在非公开内容上表现出一致的识别模式,其AUROC得分为0.82(95%置信区间:0.60–0.96),但因测试样本量小,置信区间较宽,不确定性较大。相比之下,较小的GPT-4o Mini模型对非公开内容的识别能力极弱,其AUROC为0.56(0.28–0.83)。通过固定截止日期测试多个模型,部分控制了语言随时间变化带来的偏差,但模型规模、架构及训练数据构成差异仍限制了结论强度。该初步结果强调了提升企业训练数据来源透明度及建立正式人工智能内容训练授权机制的重要性。本文主要贡献在于区分公共与非公共数据进行独立分析。

原文摘要 · Abstract (English)

Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate whether OpenAI's large language models show recognition of copyrighted content. Our results based on this small sample suggest that GPT-4o, OpenAI's more recent and capable model, exhibits patterns consistent with recognition of pay-walled book content, with an AUROC score of 0.82 (95% bootstrapped CI: 0.60-0.96), though this wide confidence interval reflects substantial uncertainty due to the limited number of books tested. GPT-4o Mini, as a much smaller model, shows little recognition of any O'Reilly Media content with an AUROC score of 0.56 (0.28-0.83) for non-public data. Testing multiple models, with the same cutoff date, provides a partial control for potential language shifts over time that might bias our findings, though differences in model size, architecture, and potentially training data composition limit the strength of this control. These preliminary results underscore the importance of increased corporate transparency regarding pre-training data sources and the development of formal licensing frameworks for AI content training. Our principal contribution is our examination of public and non public data separately.

大模型安全版权识别数据隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。