arXiv:2412.03107cs.AI2024-12被引 5

为大模型设计可验证的多比特水印,解决身份识别与隐私保护难题

CredID: Credible Multi-Bit Watermark for Large Language Models Identification

  • 引入可信第三方协调多方厂商嵌入水印,保护用户提示隐私
  • 提出新型多比特水印算法,实现高精度识别且不影响文本质量
  • 开源工具包支持研究,适合需跨厂商身份验证的场景

大型语言模型在复杂自然语言任务中广泛应用,但因缺乏身份识别机制而引发隐私与安全问题。本文提出一种多方可信水印框架(CredID),通过可信第三方(TTP)协调多个模型厂商,在不传输用户提示的前提下生成水印文本,并在提取阶段由TTP协同各厂商完成水印验证,确保水印可信性的同时保护厂商隐私。现有水印算法在文本质量、信息容量和鲁棒性方面表现不足,难以满足多样化识别需求。为此,我们提出一种新型多比特水印算法并发布开源工具包以促进研究。实验表明,该方案在不降低文本质量的前提下显著提升水印可信度与效率,并成功实现了对多个大模型厂商的高精度识别。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are widely used in complex natural language processing tasks but raise privacy and security concerns due to the lack of identity recognition. This paper proposes a multi-party credible watermarking framework (CredID) involving a trusted third party (TTP) and multiple LLM vendors to address these issues. In the watermark embedding stage, vendors request a seed from the TTP to generate watermarked text without sending the user's prompt. In the extraction stage, the TTP coordinates each vendor to extract and verify the watermark from the text. This provides a credible watermarking scheme while preserving vendor privacy. Furthermore, current watermarking algorithms struggle with text quality, information capacity, and robustness, making it challenging to meet the diverse identification needs of LLMs. Thus, we propose a novel multi-bit watermarking algorithm and an open-source toolkit to facilitate research. Experiments show our CredID enhances watermark credibility and efficiency without compromising text quality. Additionally, we successfully utilized this framework to achieve highly accurate identification among multiple LLM vendors.

大模型水印可信识别隐私保护多厂商协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。