arXiv:2607.16010cs.CRcs.CL2026-07

实测三大文本水印均无法通过司法取证检验,全数失效或误判。

AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation

  • 用法律取证标准测试三种主流水印方法,聚焦语义不变改写攻击
  • 水印100%被改写消除,虚假未检测率高达70%~83%,误标人类文本超5%
  • 水印系统根本无法满足法庭证据要求,适合关注AI监管与司法可信性的研究者

各国正强制要求大模型生成内容添加水印。欧盟《人工智能法案》要求标记“足够可靠且鲁棒”,加州SB 942则要求披露“永久或极难移除”。两者均基于未经验证的假设:水印检测可作为法庭证据。本文直接检验该假设。我们评估三种代表性水印方法——KGW、Unigram及SynthID-Text(MarkLLM实现)——是否符合达伯特可采性标准与NIST SP 800-86数字取证流程。提出“司法取证准备度评分”(FRS)框架,包含12项标准、3个必过关卡与60分制。重点考察语义保持型改写攻击,因其在法律上合理且难以被认定为篡改。结果显示严重证据效力问题:在15种提示下共846次有效改写中,所有初始检测到的KGW与Unigram文本均失去水印——条件去除率达100%;SynthID仅略好,为98.3%。攻击前假阴性率已高:KGW 70%,Unigram 83%,SynthID 80%。SynthID还错误标记5.4%的人类改写文本为AI生成,出现18.6%悖论率,其原始水印输出中有80%落入不确定区间。三方法均未满足五个达伯特因素中的超过两个。尽管FRS评分体系运行正常,但无法完全揭示其取证无用性,未来框架设计需注意此局限。测试结果表明,当前配置无法达到法院对证据的要求。

原文摘要 · Abstract (English)

Governments are increasingly mandating that LLM-generated content carry watermarks. The EU AI Act calls for markings that are "sufficiently reliable and robust." California's SB 942 requires disclosure that is "permanent or extraordinarily difficult to remove." Both mandates rest on an untested assumption: that watermark detection yields evidence reliable enough for courts. This paper tests that assumption directly. We evaluate three representative LLM watermarking methods -- KGW, Unigram, and the MarkLLM implementation of SynthID-Text -- against the Daubert admissibility criteria and the NIST SP 800-86 digital forensic process. To structure this evaluation, we propose a Forensic Readiness Score (FRS) framework with 12 criteria, three mandatory gates, and a 60-point scoring system. We focus on meaning-preserving paraphrase as the attack vector, since it is both legally realistic and difficult to dismiss as evidence tampering. The results raise serious evidentiary concerns. Out of 846 valid paraphrase runs across 15 diverse prompts per method, every single initially-detected KGW and Unigram text lost its watermark after paraphrasing -- 100% conditional removal. SynthID fared only slightly better at 98.3%. Even before any attack, false-negative rates were already high: 70% for KGW, 83% for Unigram, 80% for SynthID. The SynthID configuration also flagged 5.4% of paraphrased human-written controls as AI-generated and showed an 18.6% paradox rate, with 80% of its own pristine watermarked output landing in the uncertainty deadband. None of the three methods satisfy more than two of five Daubert factors. We also find that the FRS point-based scoring system, despite working as designed, cannot fully capture forensic uselessness -- a limitation worth noting for future framework design. These configurations, as tested, do not meet the evidentiary bar that courts require.

AI水印司法取证大模型安全证据可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。