提出De-mark框架,可有效移除大模型中的n-gram水印
De-mark: Watermark Removal in Large Language Models
- 采用随机探测策略识别水印的红绿列表
- 在Llama3和ChatGPT上实现高效水印移除
- 适合研究生成内容溯源与反水印技术者
水印技术通过在语言模型生成内容中嵌入隐蔽信息,为机器生成内容提供识别途径。然而,现有水印方案的鲁棒性尚未充分探索。本文提出De-mark框架,可有效移除基于n-gram的水印。该方法采用新型查询策略——随机选择探测,用于评估水印强度并定位n-gram水印中的红绿列表。在Llama3和ChatGPT等主流语言模型上的实验表明,De-mark在水印移除与利用任务中均表现高效且有效。
原文摘要 · Abstract (English)
Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models (LMs). However, the robustness of the watermarking schemes has not been well explored. In this paper, we present De-mark, an advanced framework designed to remove n-gram-based watermarks effectively. Our method utilizes a novel querying strategy, termed random selection probing, which aids in assessing the strength of the watermark and identifying the red-green list within the n-gram watermark. Experiments on popular LMs, such as Llama3 and ChatGPT, demonstrate the efficiency and effectiveness of De-mark in watermark removal and exploitation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。