arXiv:2512.13325cs.CRcs.AI2025-12中稿 · and presented at t…

测试发现大模型能检测文本水印,但无法提取密钥。

Security and Detectability Analysis of Unicode Text Watermarking Methods against Large Language Models

  • 用三种实验对比10种Unicode水印方法在6个大模型上的表现。
  • 最新推理模型可识别水印存在,但无法恢复隐藏信息。
  • 适合关注文本安全与对抗生成内容的研究者阅读。

随着大语言模型的广泛应用,数字文本安全日益重要。用户担心数据被用于训练模型或难以区分机器生成文本与人工写作。数字水印通过在数据中嵌入不可见标记提供额外保护。然而,现有水印方法是否对大语言模型具备安全性与隐蔽性仍缺乏系统研究。本文在三个受控实验中,实现了十种现有的Unicode文本水印方法,并在六款大语言模型(GPT-5、GPT-4o、Teuken 7B、Llama 3.3、Claude Sonnet 4、Gemini 2.5 Pro)上进行评估。结果表明,尤其是最新的推理型模型能够检测到水印的存在;但所有模型均无法提取水印,除非获得源代码实现细节。研究讨论了对安全研究人员与实践者的启示,并指出了未来需解决的安全挑战方向。

原文摘要 · Abstract (English)

Securing digital text is becoming increasingly relevant due to the widespread use of large language models. Individuals' fear of losing control over data when it is being used to train such machine learning models or when distinguishing model-generated output from text written by humans. Digital watermarking provides additional protection by embedding an invisible watermark within the data that requires protection. However, little work has been taken to analyze and verify if existing digital text watermarking methods are secure and undetectable by large language models. In this paper, we investigate the security-related area of watermarking and machine learning models for text data. In a controlled testbed of three experiments, ten existing Unicode text watermarking methods were implemented and analyzed across six large language models: GPT-5, GPT-4o, Teuken 7B, Llama 3.3, Claude Sonnet 4, and Gemini 2.5 Pro. The findings of our experiments indicate that, especially the latest reasoning models, can detect a watermarked text. Nevertheless, all models fail to extract the watermark unless implementation details in the form of source code are provided. We discuss the implications for security researchers and practitioners and outline future research opportunities to address security concerns.

文本安全水印检测大模型防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。