arXiv:2512.07038cs.CRcs.LG2025-12被引 1

为语言模型生成可信赖的溯源与水印机制提供理论框架

Ideal Attribution and Faithful Watermarks for Language Models

  • 提出'账本'机制记录用户与模型交互历史,实现确定性溯源
  • 将水印设计目标明确为忠实反映理想溯源机制,提升可解释性
  • 为未来水印方案提供可衡量的理想标准,适合研究者参考

我们引入理想溯源机制,一种用于推理字符串溯源决策的形式化抽象。其核心是账本——一个追加式的日志,记录模型与用户之间的提示-响应交互历史。每种机制基于账本和明确的选择准则,产生确定性决策,适合作为溯源的基准真值。我们将水印方案的设计目标定义为对理想溯源机制的忠实表示。这一新视角带来概念清晰性,用统一语言替代零散的概率陈述,准确描述各方案的保证。它还支持对未来水印方案所需特性的精确推理,即使当前尚无实现。该框架提供了路线图,明确了在理想设定下可实现的保障,并指导实际应用中值得追求的目标。

原文摘要 · Abstract (English)

We introduce ideal attribution mechanisms, a formal abstraction for reasoning about attribution decisions over strings. At the core of this abstraction lies the ledger, an append-only log of the prompt-response interaction history between a model and its user. Each mechanism produces deterministic decisions based on the ledger and an explicit selection criterion, making it well-suited to serve as a ground truth for attribution. We frame the design goal of watermarking schemes as faithful representation of ideal attribution mechanisms. This novel perspective brings conceptual clarity, replacing piecemeal probabilistic statements with a unified language for stating the guarantees of each scheme. It also enables precise reasoning about desiderata for future watermarking schemes, even when no current construction achieves them, since the ideal functionalities are specified first. In this way, the framework provides a roadmap that clarifies which guarantees are attainable in an idealized setting and worth pursuing in practice.

溯源机制水印设计语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。