arXiv:2602.05252cs.CL2026-02ACL被引 2

首个可交互的LLM版权风险检测系统,助力发现模型泄露内容

Copyright Detective: A Forensic System to Evidence LLMs Flickering Copyright Leakage Risks

  • 将版权问题视为证据发现过程,融合多种检测方法
  • 支持逐字记忆与改写泄露的系统性审计
  • 适合评估大模型版权风险,尤其适用于黑盒场景

我们提出Copyright Detective,首个用于检测、分析和可视化LLM输出潜在版权风险的交互式取证系统。由于版权法的复杂性,该系统将侵权与合规判断视为证据发现过程而非静态分类任务。它在统一可扩展框架中集成内容回忆测试、改写级相似性分析、说服性越狱探测及去学习验证等多种检测范式。通过交互式提示、响应收集与迭代工作流,系统能够系统性审计逐字记忆与改写级内容泄露,支持在仅具备黑盒访问权限的情况下,实现大模型版权风险的负责任部署与透明评估。

原文摘要 · Abstract (English)

We present Copyright Detective, the first interactive forensic system for detecting, analyzing, and visualizing potential copyright risks in LLM outputs. The system treats copyright infringement versus compliance as an evidence discovery process rather than a static classification task due to the complex nature of copyright law. It integrates multiple detection paradigms, including content recall testing, paraphrase-level similarity analysis, persuasive jailbreak probing, and unlearning verification, within a unified and extensible framework. Through interactive prompting, response collection, and iterative workflows, our system enables systematic auditing of verbatim memorization and paraphrase-level leakage, supporting responsible deployment and transparent evaluation of LLM copyright risks even with black-box access.

版权风险LLM审计取证系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。