通过检测令牌重复概率,零样本识别大模型生成文本
Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability

- 提出新指标Telescope Perplexity,衡量文本中令牌重复的异常性
- 在预训练早期即能捕捉到大模型生成文本的特征,零样本检测效果优异
- 无需微调,适用于多种数据集和扰动场景,效率高于现有方法
区分大语言模型(LLM)生成文本与人类写作是一项关键而艰巨的挑战。尽管LLM经过训练以模仿人类写作风格,但其早期训练过程中形成的对令牌重复的强烈规避倾向会残留为一种‘遗迹启发式’(Vestigial Heuristic),在生成文本中被激活,成为与人类写作的关键差异。为此,我们提出Telescope Perplexity,一种评估模型条件概率 $P(s_i | s_{1:i})$ 的指标,用以探测文本中的令牌重复模式。实证研究显示,Telescope Perplexity在预训练初期即已显现,且在多种数据集(包括我们引入的现代评测集)、参考模型和扰动方案下均实现领先或相当的零样本检测性能,同时比其他方法更具效率。
原文摘要 · Abstract (English)
Distinguishing Large Language Model (LLM) generated text from human writing is a critical and difficult challenge. While LLMs are trained to write like humans, we hypothesize that this training leaves an indelible mark. LLMs develop a particularly strong aversion to token repetition very early in training. This bias persists as a ''Vestigial Heuristic'' (a developmental artifact) that is activated in LLM-generated text, separating LLM from human writing. To probe this phenomenon, we introduce Telescope Perplexity, a metric that evaluates the token repetition of the model, $P(s_i | s_{1:i})$ . Our empirical investigation reveals that the Telescope Perplexity signature emerges early in pre-training, and Telescope Perplexity empirically enables highly effective zero-shot LLM detection. We show state-of-the-art or competitive performance across diverse datasets (including modern evaluation sets we introduce), reference models, and perturbation schemes with greater efficiency than other methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。