提升大模型水印抗编辑能力,实现低误报率的多比特鲁棒标记。
CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking

- 基于固定命中率校准水印信道,生成逐标记的对数似然比,支持软解码。
- 在令牌编辑和改写下仍保持低误报率,优于现有多比特水印方法。
- 适合需要可信生成溯源且对抗篡改的场景,如内容审计与版权保护。
大模型输出的可靠溯源需要在编辑后仍保持鲁棒性的多比特水印,并严格控制误报率。现有基于纠错码(ECC)的水印方法主要依赖硬判决解码,忽略了令牌级别的可靠性信息。本文提出CORE-BREW,作为块级BREW的恒定命中率扩展,通过设定目标命中率p-star校准水印信道,获得闭式表达的每标记对数似然比(LLRs),支持有原则的软解码。该方法支持两种检测模式:严格安全模式(Strict-Safe)保留预定码字接受区域的有界距离特性;FPR校准模式(FPR-Calibrated)采用基于似然的评分与轻量级列表解码,刻画误报率(FPR)与真正率(TPR)的权衡关系。在开源大模型上进行令牌级编辑与重述实验表明,相比先前多比特水印基线,CORE-BREW在保持相近语义质量的同时,显著提升了低误报率下的区分能力和鲁棒性。
原文摘要 · Abstract (English)
Reliable provenance for LLM outputs requires multi-bit watermarks that remain robust under editing while maintaining strict false-positive control. Existing ECC-based LLM watermarks rely largely on hard-decision decoding, discarding token-level reliability information. We propose CORE-BREW, a Constant-hit-Rate Embedding extension of block-wise BREW for robust multi-bit watermarking. CORE-BREW calibrates the watermark channel by targeting a fixed hit rate p-star, yielding closed-form per-token log-likelihood ratios (LLRs) for principled soft-decision decoding. It supports two detection modes: Strict-Safe, which preserves the bounded-distance designated-codeword acceptance region, and FPR-Calibrated, which uses likelihood-based scoring and lightweight list decoding to characterize the FPR-TPR trade-off. Experiments on open-source LLMs under token-level edits and paraphrasing demonstrate improved low-FPR discrimination and robustness over prior multi-bit watermarking baselines while maintaining comparable semantic quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。