arXiv:2505.12655cs.CRcs.AI2025-05EMNLP被引 5

用大模型理解能力反制其非法实时抓取网页内容,保护原创者权益

Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language Models

  • 利用大模型自身语义理解能力构建防御机制
  • 实测将防御成功率从2.5%提升至88.6%
  • 适合关注网络版权保护的内容创作者

网络知识产权(IP)如网页内容的保护日益成为关键问题。具备在线检索能力的大语言模型(LLMs)虽方便信息获取,却常侵犯原创内容作者权利。用户越来越依赖大模型生成结果,导致对原始信息源的直接访问减少,严重削弱内容创作者的创作动力,最终使网络空间充斥大量AI生成内容。为此,我们提出一种新型防御框架,利用大模型自身的语义理解能力,帮助网页内容创作者防范未经授权的大模型实时提取与再分发。该方法基于合理设计,有效解决难以处理的黑箱优化问题。真实世界实验表明,我们的方法在不同大模型上将防御成功率从2.5%提升至88.6%,显著优于传统的配置限制类防御手段。

原文摘要 · Abstract (English)

The protection of cyber Intellectual Property (IP) such as web content is an increasingly critical concern. The rise of large language models (LLMs) with online retrieval capabilities enables convenient access to information but often undermines the rights of original content creators. As users increasingly rely on LLM-generated responses, they gradually diminish direct engagement with original information sources, which will significantly reduce the incentives for IP creators to contribute, and lead to a saturating cyberspace with more AI-generated content. In response, we propose a novel defense framework that empowers web content creators to safeguard their web-based IP from unauthorized LLM real-time extraction and redistribution by leveraging the semantic understanding capability of LLMs themselves. Our method follows principled motivations and effectively addresses an intractable black-box optimization problem. Real-world experiments demonstrated that our methods improve defense success rates from 2.5% to 88.6% on different LLMs, outperforming traditional defenses such as configuration-based restrictions.

知识产权大模型防御内容保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。