arXiv:2605.13095cs.CRcs.AI2026-05被引 1

水印应视为监控机制,而非仅防伪造工具。

Watermarking Should Be Treated as a Monitoring Primitive

论文配图:Watermarking Should Be Treated as a Monitoring Primitive
图 1 · 摘自论文原文
  • 提出观察者威胁模型,可聚合水印信号推断实体信息
  • 零比特水印在多密钥下仍能实现归属追踪
  • 适合关注生成模型安全监控与隐私风险的研究者

水印被广泛用于生成模型的来源追溯、责任归属和安全监控,但现有评估通常只考虑单样本级的逃避检测或误报攻击。我们认为水印应被视为一种监控原语,并指出在每个实体有专属密钥和消息、且检测器可访问的前提下,内部监控不可避免。本文引入基于观察者的威胁模型,展示观察者可通过聚合多个输出的水印信号推断实体级信息,证明即使零比特水印在多密钥设置下仍可实现归属。此外,外部监控可能随时间通过持久存在的密钥相关统计结构自然出现,但这一现象取决于水印设计,可通过保持分布或不可检测方案缓解。研究揭示了归属与监控之间的根本双重用途矛盾,呼吁超越单样本鲁棒性,评估水印在聚合与观察者能力下的表现。

原文摘要 · Abstract (English)

Watermarking is widely proposed for provenance, attribution, and safety monitoring in generative models, yet is typically evaluated only under adversaries who attempt to evade detection or induce false positives at the level of individual samples. We argue that watermarking should be treated as a monitoring primitive, and that internal monitoring is unavoidable given per-entity attribution keys and messages, as well as detector access. We introduce an observer-based threat model in which observers can aggregate watermark signals across outputs to infer entity-level information, showing that even zero-bit watermarking enables attribution under multi-key settings. We further show that external monitoring can emerge over time from persistent, key-dependent statistical structure, although this depends on watermark design and may be mitigated by distribution-preserving or undetectable schemes. Our findings reveal a fundamental dual-use tension between attribution and monitoring, motivating evaluation of watermarking beyond per-sample robustness to account for aggregation and observer-based capabilities.

水印监控生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。