arXiv:2603.14968cs.CRcs.CL2026-03ACL

提出无需密钥的第三方水印检测框架,实现对大模型生成内容的独立审计。

Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework

  • 通过代理模型放大水印信号,将检测转为相对假设检验。
  • 在多个模型和攻击场景下均保持高检测准确率与鲁棒性。
  • 适合监管机构、第三方审计或需要独立验证水印的场景。

尽管水印是保障大模型生成内容溯源的关键机制,现有密钥方案将检测与注入强绑定,需访问密钥或服务商专属检测器才能验证,导致独立审计难以实现,且可能威胁模型安全或依赖服务提供商的不透明声明。为解决这一困境,我们提出 TTP-Detect——首个黑盒非侵入式第三方水印验证框架。该框架解耦检测与注入,将验证重构为相对假设检验问题。通过代理模型增强水印相关信号,并结合多种互补的相对度量,评估查询文本与水印分布的一致性。在代表性水印方案、数据集和模型上的大量实验表明,TTP-Detect 在多种攻击下仍具备优异的检测性能与鲁棒性。

原文摘要 · Abstract (English)

While watermarking serves as a critical mechanism for LLM provenance, existing secret-key schemes tightly couple detection with injection, requiring access to keys or provider-side scheme-specific detectors for verification. This dependency creates a fundamental barrier for real-world governance, as independent auditing becomes impossible without compromising model security or relying on the opaque claims of service providers. To resolve this dilemma, we introduce TTP-Detect, a pioneering black-box framework designed for non-intrusive, third-party watermark verification. By decoupling detection from injection, TTP-Detect reframes verification as a relative hypothesis testing problem. It employs a proxy model to amplify watermark-relevant signals and a suite of complementary relative measurements to assess the alignment of the query text with watermarked distributions. Extensive experiments across representative watermarking schemes, datasets and models demonstrate that TTP-Detect achieves superior detection performance and robustness against diverse attacks.

水印检测大模型黑盒审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。