给智能体系统加数字水印,防抄袭还能不影响性能。
On Protecting Agentic Systems' Intellectual Property via Watermarking
- 通过改变等效动作路径分布来嵌入水印信号。
- 在三个复杂场景中检测准确率高,性能损失可忽略。
- 适合保护自主推理智能体的知识产权,防自适应攻击。
大型语言模型演变为具备自主推理与工具调用能力的智能体系统,创造了巨大的知识产权价值。我们证明这类系统极易遭受模仿攻击:攻击者可通过训练模仿模型来窃取其专有功能。关键问题是,现有LLM水印技术在此场景失效,因为真实智能体系统常以灰盒形式运行,内部推理过程无法验证。本文提出AGENTWM,首个专为智能体模型设计的水印框架。该框架利用动作序列的语义等价性,通过轻微偏移功能等效的工具执行路径分布来嵌入水印,使可验证信号直接存在于可见动作轨迹中,对用户不可察觉。我们构建了自动化水印生成管道和严格的统计假设检验流程用于验证。在三个复杂领域上的大量实验表明,AGENTWM在几乎不影响智能体性能的前提下实现了高检测准确率。结果证实,该方法能有效抵御自适应攻击——攻击者若想移除水印,必将严重损害被窃模型的可用性。
原文摘要 · Abstract (English)
The evolution of Large Language Models (LLMs) into agentic systems that perform autonomous reasoning and tool use has created significant intellectual property (IP) value. We demonstrate that these systems are highly vulnerable to imitation attacks, where adversaries steal proprietary capabilities by training imitation models on victim outputs. Crucially, existing LLM watermarking techniques fail in this domain because real-world agentic systems often operate as grey boxes, concealing the internal reasoning traces required for verification. This paper presents AGENTWM, the first watermarking framework designed specifically for agentic models. AGENTWM exploits the semantic equivalence of action sequences, injecting watermarks by subtly biasing the distribution of functionally identical tool execution paths. This mechanism allows AGENTWM to embed verifiable signals directly into the visible action trajectory while remaining indistinguishable to users. We develop an automated pipeline to generate robust watermark schemes and a rigorous statistical hypothesis testing procedure for verification. Extensive evaluations across three complex domains demonstrate that AGENTWM achieves high detection accuracy with negligible impact on agent performance. Our results confirm that AGENTWM effectively protects agentic IP against adaptive adversaries, who cannot remove the watermarks without severely degrading the stolen model's utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。