arXiv:2505.06827cs.CRcs.AI2025-05ACL被引 3

实验证明水印在真实场景中比理论预测更难消除。

Sandcastles in the Storm: Revisiting the (Im)possibility of Strong Watermarking

  • 通过大规模实验发现水印在数百次编辑后仍可保留
  • 自动攻击仅26%成功率,人类评估降至10%
  • 适合关注文本生成安全与水印设计的研究者

文本水印对防止滥用至关重要。然而近期理论认为,通过随机游走攻击可在不降低质量的前提下擦除水印。该结论依赖两个假设:快速混合(水印迅速消失)和可靠质量控制(检测器精准指导修改)。我们通过大规模实验和人工评估发现,混合过程实际缓慢:100%的扰动文本在数百次编辑后仍保留原始痕迹;同时,最先进的质量检测器准确率仅为77%,导致攻击中错误累积。最终,自动化攻击仅26%成功率,人类评审下进一步降至10%。这些结果挑战了水印必然被破坏的论断,表明现实中的慢混合与质量控制缺陷使水印远比理论模型所预测的更鲁棒。理想化攻击与实际可行性间的差距,凸显需发展更强水印方法及更真实的攻击评估框架。

原文摘要 · Abstract (English)

Watermarking AI-generated text is critical for combating misuse. Yet recent theoretical work argues that any watermark can be erased via random walk attacks that perturb text while preserving quality. However, such attacks rely on two key assumptions: (1) rapid mixing (watermarks dissolve quickly under perturbations) and (2) reliable quality preservation (automated quality oracles perfectly guide edits). Through large-scale experiments and human-validated assessments, we find mixing is slow: 100% of perturbed texts retain traces of their origin after hundreds of edits, defying rapid mixing. Oracles falter, as state-of-the-art quality detectors misjudge edits (77% accuracy), compounding errors during attacks. Ultimately, attacks underperform: automated walks remove watermarks just 26% of the time -- dropping to 10% under human quality review. These findings challenge the inevitability of watermark removal. Instead, practical barriers -- slow mixing and imperfect quality control -- reveal watermarking to be far more robust than theoretical models suggest. The gap between idealized attacks and real-world feasibility underscores the need for stronger watermarking methods and more realistic attack models.

文本水印生成安全对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。