提出新特征检测水印,显著提升长文本和短文本的鲁棒性。
Stability-Aware Feature Design for Robust Watermark Detection in Machine-Generated Text

- 基于局部统计与稳定性动态设计新特征,融合高阶模式与自相关信号。
- 在八轮改写下仍保持87.8%以上AUC,相比基线提升10-15个百分点。
- 单模型可跨模型、跨领域通用,无需重新训练。
大型语言模型的广泛应用加剧了区分人工与机器生成文本的需求。水印技术提供了一条有前景的路径,但现有检测器在多次改写及短文本场景下性能急剧下降。本文提出模式稳定性得分(PSS),一种利用局部统计特征与改写变体间稳定性动态的新检测框架。该方法结合全局与局部z-score特征,引入运行长度模式的高阶统计量,并通过自相关信号与改写深度上的稳定性得分增强表达能力。在三个基准数据集(PG-19、CNN/DailyMail、WikiText)上,使用多个LLM(Llama-3-8B、Qwen2-7B)和改写器(Mistral-7B、Qwen2-7B、Gemma-7B)进行系统评估,压力测试覆盖最多八轮改写。相比传统z-score阈值基线与部分先进深度学习方法,本方法在不同词元长度下检测AUC提升超过10-15个百分点。此外,跨领域实验表明,单一通用分类器可泛化至不同LLM、改写器与文本领域,无需再训练,即使所有组件均异于训练环境,仍保持高于87.8%的AUC。
原文摘要 · Abstract (English)
The widespread adoption of large language models (LLMs) has intensified the demand for principled methods to distinguish human from machine-generated text. Watermarking provides a promising avenue, yet existing detectors exhibit sharp performance deterioration under multiple paraphrasing and when applied to shorter texts. We introduce Pattern Stability Score (PSS), a novel detection framework that leverages local statistical features and stability dynamics across paraphrased variants. Specifically, the proposed method combines global and local z-score features with higher-order statistics of run-length patterns, enriched by autocorrelation signals and stability scores computed over paraphrase depth. Numerical evaluations are performed on three benchmark datasets (PG-19, CNN/DailyMail, and WikiText) using multiple LLMs (Llama-3-8B, Qwen2-7B) and paraphrasers (Mistral-7B, Qwen2-7B, Gemma-7B), systematically stress-testing robustness under up to eight rounds of paraphrasing. Compared to prior z-score thresholding baselines and some state-of-the-art deep learning methods, our approach improves detection AUC (area under the receiver operating characteristic curve) by over 10-15 percentage points across different token lengths. Additionally, extensive cross-domain experiments demonstrate that a single universal classifier generalizes across different LLMs, paraphrasers, and text domains without retraining, maintaining above 87.8% AUC even when all components differ from training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。