arXiv:2409.09739cs.CRcs.CL2024-09被引 6

为每个用户生成个性化水印,保护大模型版权并追踪使用者

PersonaMark: Personalized LLM watermarking for model protection and user attribution

  • 用句子结构隐藏水印信息,不改变模型输出质量
  • 支持多用户,水印检测准确率高且不影响文本自然度
  • 适合需防抄袭和溯源的私有大模型应用场景

定制化大语言模型(LLMs)的快速发展带来了便利,但也加剧了版权与隐私保护的担忧。随着私有LLMs的广泛应用,保护模型版权和确保数据隐私至关重要。文本水印已成为检测AI生成内容、保护模型的有效手段。然而,现有方法难以实现针对每位用户的个性化水印,限制了问责与追溯能力。本文提出PersonaMark,一种新型个性化文本水印方案,旨在保护LLM版权并增强可追溯性。该方法利用句子结构作为水印信息的隐蔽载体,优化生成过程以保持输出自然性。通过个性化哈希函数,为每位用户嵌入唯一水印,实现高质量文本生成且不降低模型性能。该方法高效可扩展,支持大规模用户,采用多用户哈希机制。据我们所知,这是首个探索LLM个性化水印的研究。我们在四种LLM上进行了广泛评估,分析困惑度、情感、对齐度和可读性等指标。结果表明,PersonaMark能有效保持文本质量,实现无偏水印插入,并具备强鲁棒性水印检测能力,同时对模型行为影响极小。

原文摘要 · Abstract (English)

The rapid advancement of customized Large Language Models (LLMs) offers considerable convenience. However, it also intensifies concerns regarding the protection of copyright/confidential information. With the extensive adoption of private LLMs, safeguarding model copyright and ensuring data privacy have become critical. Text watermarking has emerged as a viable solution for detecting AI-generated content and protecting models. However, existing methods fall short in providing individualized watermarks for each user, a critical feature for enhancing accountability and traceability. In this paper, we introduce PersonaMark, a novel personalized text watermarking scheme designed to protect LLMs' copyrights and bolster accountability. PersonaMark leverages sentence structure as a subtle carrier of watermark information and optimizes the generation process to maintain the natural output of the model. By employing a personalized hashing function, unique watermarks are embedded for each user, enabling high-quality text generation without compromising the model's performance. This approach is both time-efficient and scalable, capable of handling large numbers of users through a multi-user hashing mechanism. To the best of our knowledge, this is a pioneer study to explore personalized watermarking in LLMs. We conduct extensive evaluations across four LLMs, analyzing various metrics such as perplexity, sentiment, alignment, and readability. The results validate that PersonaMark preserves text quality, ensures unbiased watermark insertion, and offers robust watermark detection capabilities, all while maintaining the model's behavior with minimal disruption.

大模型安全水印技术用户溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。