arXiv:2511.20623cs.AI2025-11

为内容创作者提供可验证模型训练数据是否侵权的开源工具。

Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development

  • 构建易用的开源平台,支持用户上传内容检测是否被模型训练使用。
  • 通过优化接口减少10%-30%计算开销,提升相似内容检测效率。
  • 适合独立创作者、研究者及关注AI伦理与版权合规的群体使用。

大型语言模型(LLM)的广泛应用引发了对其训练数据中未经授权使用受版权保护内容的严重关切。现有检测框架如DE-COP计算成本高,对独立创作者不友好。随着法律审查日益严格,亟需一种可扩展、透明且用户友好的解决方案。本文提出一个开源版权检测平台,使内容创作者能够验证其作品是否被用于LLM训练数据。该方法通过提升易用性、改进相似性检测、优化数据集验证流程,并借助高效API调用将计算开销降低10%-30%。结合直观的用户界面与可扩展后端,该框架增强了AI开发的透明度与伦理合规性,为负责任的AI发展和版权执法研究奠定基础。

原文摘要 · Abstract (English)

The widespread use of Large Language Models (LLMs) raises critical concerns regarding the unauthorized inclusion of copyrighted content in training data. Existing detection frameworks, such as DE-COP, are computationally intensive, and largely inaccessible to independent creators. As legal scrutiny increases, there is a pressing need for a scalable, transparent, and user-friendly solution. This paper introduce an open-source copyright detection platform that enables content creators to verify whether their work was used in LLM training datasets. Our approach enhances existing methodologies by facilitating ease of use, improving similarity detection, optimizing dataset validation, and reducing computational overhead by 10-30% with efficient API calls. With an intuitive user interface and scalable backend, this framework contributes to increasing transparency in AI development and ethical compliance, facilitating the foundation for further research in responsible AI development and copyright enforcement.

版权检测LLM伦理开源工具AI合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。