arXiv:2605.10977cs.CRcs.AI2026-05被引 4

提出一种语义级水印技术,有效抵御改写攻击且不降低文本质量。

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks

论文配图:PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks
图 1 · 摘自论文原文
  • 在隐空间语义聚类上构建水印,通过共享随机性与语义历史同步
  • 在多种改写攻击下仍保持高检测率,文本质量无损失
  • 适合需可靠生成内容溯源的场景,如学术写作、新闻审核

大语言模型生成文本的水印技术是检测生成内容、实现负责任部署的有力手段。然而,现有方法常易受语义不变攻击(如改写)影响。本文提出PASA,一种理论驱动、鲁棒且无失真的水印算法,在语义层面嵌入并检测水印。PASA基于隐空间中的语义聚类,利用秘密密钥和语义历史同步的共享随机性,建立标记序列与辅助序列间的分布依赖关系。该设计基于我们提出的理论框架,刻画了嵌入-检测对的联合最优性,实现了检测准确率、鲁棒性与失真之间的根本权衡。在多个大语言模型及语义不变攻击下的评估表明,即使在强改写攻击下,PASA仍保持鲁棒性,同时维持高文本质量,显著优于标准词汇空间基线。消融实验进一步验证了超参数选择的有效性。

原文摘要 · Abstract (English)

Watermarking for large language models (LLMs) is a promising approach for detecting LLM-generated text and enabling responsible deployment. However, existing watermarking methods are often vulnerable to semantic-invariant attacks, such as paraphrasing. We propose PASA, a principled, robust, and distortion-free watermarking algorithm that embeds and detects a watermark at the semantic level. PASA operates on semantic clusters in a latent embedding space and constructs a distributional dependency between token and auxiliary sequences via shared randomness synchronized by a secret key and semantic history. This design is grounded in our theoretical framework that characterizes a jointly optimal embedding-detection pair, achieving the fundamental trade-offs among detection accuracy, robustness, and distortion. Evaluations across multiple LLMs and semantic-invariant attacks demonstrate that PASA remains robust even under strong paraphrasing attacks while preserving high text quality, outperforming standard vocabulary-space baselines. Ablation studies further validate the effectiveness of our hyperparameter choices. Webpage: https://ai-kunkun.github.io/PASA_page/.

水印技术LLM安全语义不变攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。