arXiv:2605.23175cs.CRcs.CL2026-05被引 1

新水印框架在几乎不改变文本质量的前提下,能有效防止大模型被抄袭。

Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection

论文配图:Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection
图 1 · 摘自论文原文
  • 用密钥控制的替换机制,只换词不改意思
  • 98.2%检测率,文本质量评分最高
  • 适合需要保护模型版权的开发者和机构

专有大语言模型面临知识产权侵权风险,攻击者可通过收集输入输出对训练替代模型造成损失。水印可作为所有权验证手段,但现有方法常导致语义扭曲、事实错误及易受攻击。尤其在跨平台、多用户场景下,基于密钥的专用检测仍缺乏研究。为此,我们提出SAFESEAL,一种新型密钥控制水印框架,在保持模型实用性的同时实现强可检测性与鲁棒性。SAFESEAL通过密钥控制的锦标赛采样机制,仅替换语言项而不影响命名实体,确保语义一致性和事实准确性。检测方面,引入密钥控制对比检测器,联合编码文本与密钥,实现厂商专属且鲁棒的水印验证。我们推导了实用-可检测性的理论边界,并通过轻量模型、批处理与并行化显著降低延迟。大量实验表明,SAFESEAL在实用性、可检测性与鲁棒性上均优于基线,达到BERTScore 0.983、实体相似度0.963、98.2%检测率,人类评估中文本质量与内容保留得分最高,延迟与最快基线相当。为促进透明与社区发展,我们发布首个公开水印排行榜及交互式演示。

原文摘要 · Abstract (English)

Proprietary large language models (LLMs) face risks of intellectual property (IP) violation, as adversaries can replicate an LLM by collecting input-output pairs to train a surrogate model, causing financial setbacks. Watermarks offer a promising defense to verify ownership, but existing methods often struggle with semantic distortion, factual inconsistency, and adversarial attacks. In addition, key-conditioned watermarks for provider-specific detection, especially in cross-provider and multi-user scenarios, remain largely underexplored. To address these challenges, we propose SAFESEAL, a novel key-conditioned watermarking framework that achieves strong detectability with minimal impact on model utility, effectively balancing detectability, utility, and robustness. SAFESEAL preserves named entities while substituting linguistic terms with context-aware synonyms through a key-conditioned Tournament sampling mechanism, maintaining semantic fidelity and factual consistency. For detection, we introduce a key-conditioned contrastive detector that jointly encodes the text and key, enabling provider-specific and robust watermark verification. We derive theoretical bounds on the utility-detectability trade-off and significantly reduce latency through lightweight models, batching, and parallelism. Extensive experiments show that SAFESEAL outperforms baselines in utility, detectability, and robustness, achieving a BERTScore of 0.983, entity similarity of 0.963, a 98.2% detection rate, and the highest human ratings for text quality and content preservation, with latency comparable to the fastest baseline. To promote transparency and community-driven progress, we release the first public watermark leaderboard and an interactive demo.

水印技术模型保护大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。