arXiv:2505.01484cs.CRcs.LG2025-05被引 1
提出两种不可检测且不可移除的LLM文本水印方案,有效区分机器生成与人工内容。
LLM Watermarking Using Mixtures and Statistical-to-Computational Gaps
- 基于混合模型设计不可检测水印,隐蔽性强
- 在开放场景下实现抗篡改水印,攻击者难以移除
- 适合内容安全审核与版权保护场景
针对文本是否由大语言模型(LLM)生成的问题,现有研究普遍采用水印技术。本文提出一种闭合设置下的不可检测、基础性水印方案;在更难的开放设置中,即攻击者可访问大部分模型参数时,提出一种不可移除水印方案。该方法利用混合模型结构与统计-计算间隙机制,在保持水印隐蔽性的同时,确保其在对抗攻击下依然有效,为验证生成内容来源提供了可靠手段。
原文摘要 · Abstract (English)
Given a text, can we determine whether it was generated by a large language model (LLM) or by a human? A widely studied approach to this problem is watermarking. We propose an undetectable and elementary watermarking scheme in the closed setting. Also, in the harder open setting, where the adversary has access to most of the model, we propose an unremovable watermarking scheme.
水印LLM安全生成内容检测
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。