用推理触发检测技术,防止多智能体系统泄露版权内容。
CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems
- 在推理链中嵌入触发查询,监控中间思考过程。
- 实验显示能有效发现内容泄露,且不影响任务表现。
- 适合关注AI版权保护的研究者与开发者。
随着大语言模型演变为具备协作推理与任务执行能力的自主智能体,多智能体LLM系统已成为解决复杂问题的强大范式。然而,这类系统给版权保护带来了新挑战,尤其当敏感或受版权保护的内容通过智能体间的沟通与推理被意外复现时。现有保护技术主要聚焦于最终输出的内容检测,忽略了智能体内部更丰富、更具揭示性的推理过程。本文提出CoTGuard,一种基于推理链(CoT)触发的版权保护新框架。具体而言,通过在智能体提示中嵌入特定触发查询,可激活并监控特定推理片段,从而检测未经授权的内容复制。该方法实现了对协作智能体场景下版权违规行为的细粒度、可解释性检测。我们在多个基准上进行了广泛实验,结果表明,CoTGuard能有效发现内容泄露,且对任务性能干扰极小。研究提示,推理级监控为保障基于LLM的智能体系统中的知识产权提供了有前景的方向。
原文摘要 · Abstract (English)
As large language models (LLMs) evolve into autonomous agents capable of collaborative reasoning and task execution, multi-agent LLM systems have emerged as a powerful paradigm for solving complex problems. However, these systems pose new challenges for copyright protection, particularly when sensitive or copyrighted content is inadvertently recalled through inter-agent communication and reasoning. Existing protection techniques primarily focus on detecting content in final outputs, overlooking the richer, more revealing reasoning processes within the agents themselves. In this paper, we introduce CoTGuard, a novel framework for copyright protection that leverages trigger-based detection within Chain-of-Thought (CoT) reasoning. Specifically, we can activate specific CoT segments and monitor intermediate reasoning steps for unauthorized content reproduction by embedding specific trigger queries into agent prompts. This approach enables fine-grained, interpretable detection of copyright violations in collaborative agent scenarios. We evaluate CoTGuard on various benchmarks in extensive experiments and show that it effectively uncovers content leakage with minimal interference to task performance. Our findings suggest that reasoning-level monitoring offers a promising direction for safeguarding intellectual property in LLM-based agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。