用双语义嵌入提升大模型文本水印抗改写与翻译能力
Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings
- 融合上下文与词元级嵌入生成鲁棒水印信号
- 改写后仍可检测,翻译后依然有效,水印衰减平缓
- 适合需要内容溯源的AI安全与版权保护场景
本文提出双嵌入水印(DEW),一种基于上下文与词元级嵌入的大语言模型文本水印方案,旨在增强对改写和翻译的鲁棒性。DEW采用信号处理方法,通过代数向量空间运算对嵌入向量进行操作,生成在语义变化下渐进退化的水印信号。水印通过秘密密钥生成的伪随机矩阵对嵌入向量进行投影以实现混淆。基于底层代数分布的统计测试与基准评估被用于验证效果。多模型实验表明,DEW在保持良好文本质量的同时显著提升改写后检测率,且在翻译后仍可检测,而传统语义水印已严重失效。该结果表明DEW是负责任AI部署中保障生成文本可追溯性的实用且稳健方案。
原文摘要 · Abstract (English)
This work presents Dual-Embedding Watermarking (DEW), a semantic watermarking scheme for large language models (LLMs) that leverages contextual and token-level embeddings to enhance robustness against paraphrasing and translation. DEW utilizes a signal-processing methodology, applying algebraic vector-space operations to token and context embeddings to derive a watermark signal that degrades gracefully under semantic shifts. The method obfuscates the watermark by projecting embedding vectors through pseudo-random matrices seeded with a secret key. Relevant distributions derived from the underlying algebra are evaluated and employed for statistical testing and benchmarking of DEW. Experimental results across multiple LLMs indicate that DEW improves post-paraphrase detection while maintaining competitive text quality, and remains detectable after translation, even when prior semantic watermarks degrade significantly. These findings position DEW as a practical and robust solution for safeguarding LLM-generated text and addressing critical issues in responsible AI deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。