arXiv:2604.21223cs.CLcs.AI2026-04NeurIPS被引 5

无需训练即可检测大模型生成文本,靠隐式奖励模型实现高精度识别。

Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model

  • 利用公开模型的隐式奖励机制,不依赖标注数据或微调。
  • 在DetectRL基准上超越现有零样本与有监督方法性能。
  • 适合需要快速部署、无标注数据场景下的文本真实性检测。

大语言模型在各类任务中表现出色,但其生成类人文本的能力引发滥用担忧,亟需可靠的检测手段。本文提出一种名为IRM的新零样本方法,基于隐式奖励模型检测大模型生成文本。该模型可从公开的指令微调和基础模型中直接推导,无需偏好数据收集或额外训练。相较于依赖偏好构建与任务微调的已有方法,IRM具备更高效率与泛化性。在DetectRL基准上的实验表明,IRM在检测大模型生成文本方面表现优异,优于现有零样本及有监督方法。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their ability to generate human-like text has raised concerns about potential misuse. This underscores the need for reliable and effective methods to detect LLM-generated text. In this paper, we propose IRM, a novel zero-shot approach that leverages Implicit Reward Models for LLM-generated text detection. Such implicit reward models can be derived from publicly available instruction-tuned and base models. Previous reward-based method relies on preference construction and task-specific fine-tuning. In comparison, IRM requires neither preference collection nor additional training. We evaluate IRM on the DetectRL benchmark and demonstrate that IRM can achieve superior detection performance, outperforms existing zero-shot and supervised methods in LLM-generated text detection.

文本检测零样本大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。