提出新水印算法SeqMark,解决低熵生成任务中的水印难题。
Semantic Differentiation for Tackling Challenges in Watermarking Low-Entropy Constrained Generation Outputs
- 基于语义区分的序列级水印方法,优化输出质量与可检测性平衡。
- 在机器翻译等任务上水印检测F1提升最高达28%,保持高质量生成。
- 适用于代码生成、摘要等低熵约束生成场景,适合需内容安全的AI应用。
现有语言模型水印方法在开放式生成中有效,但在低熵输出空间的约束生成任务中表现不足。为此,我们提出SeqMark,一种具有语义区分能力的序列级水印算法,在保持生成质量的同时提升水印可检测性与隐蔽性。该方法克服了传统令牌级水印导致的序列级熵利用不足问题。此外,我们识别并改进了先前序列级水印方法中存在的区域坍塌问题——即伪随机划分语义空间使高概率输出集中于有效或无效区域,造成生成质量与水印效果的权衡。SeqMark通过区分高概率输出子空间并均匀划分有效与无效区域,确保高质量输出在各区域间均衡分布。在机器翻译、抽象摘要和代码生成等任务上,SeqMark显著提升水印检测准确率(F1最高提升28%),同时维持高生成质量。
原文摘要 · Abstract (English)
We demonstrate that while the current approaches for language model watermarking are effective for open-ended generation, they are inadequate at watermarking LM outputs for constrained generation tasks with low-entropy output spaces. Therefore, we devise SeqMark, a sequence-level watermarking algorithm with semantic differentiation that balances output quality, watermark detectability, and imperceptibility. It improves on the shortcomings of token-level watermarking algorithms that cause under-utilization of the sequence-level entropy available for constrained generation tasks. Moreover, we identify and improve upon the problem of region collapse, a different failure mode associated with prior sequence-level watermarking algorithms. This occurs because the pseudorandom partitioning of semantic space for watermarking in these approaches causes all high-probability outputs to collapse into either invalid or valid regions, leading to a trade-off in output quality and watermarking effectiveness. Instead, SeqMark differentiates the high-probable output subspace and partitions it into valid and invalid regions, ensuring the even spread of high-quality outputs among all the regions. On various constrained generation tasks like machine translation, abstractive summarization, and code generation, SeqMark substantially improves watermark detection accuracy (up to 28% increase in F1) while maintaining high generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。