arXiv:2504.12108cs.CL2025-04被引 2

用熵阈值提升大模型生成文本的可追踪性与质量

Entropy-Guided Watermarking for LLMs: A Test-Time Framework for Robust and Traceable Text Generation

  • 引入累积水印熵阈值,动态控制水印强度
  • 在MATH和GSM8K上检测率提升超80%
  • 适合需防滥用且保文本质量的生成场景

大型语言模型的快速发展引发了内容可追溯性和潜在滥用的担忧。现有采样文本水印方案常在文本质量与抗攻击检测能力间存在权衡。为此,我们提出一种新型水印方案,通过引入累积水印熵阈值,显著提升检测能力与文本质量。该方法兼容并泛化现有采样函数,增强适应性。在多个LLM上的实验表明,本方案相比现有方法大幅提升性能,在MATH和GSM8K等常用数据集上实现超过80%的改进,同时保持高检测准确率。

原文摘要 · Abstract (English)

The rapid development of Large Language Models (LLMs) has intensified concerns about content traceability and potential misuse. Existing watermarking schemes for sampled text often face trade-offs between maintaining text quality and ensuring robust detection against various attacks. To address these issues, we propose a novel watermarking scheme that improves both detectability and text quality by introducing a cumulative watermark entropy threshold. Our approach is compatible with and generalizes existing sampling functions, enhancing adaptability. Experimental results across multiple LLMs show that our scheme significantly outperforms existing methods, achieving over 80\% improvements on widely-used datasets, e.g., MATH and GSM8K, while maintaining high detection accuracy.

大模型水印文本追踪生成安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。