arXiv:2601.06586cs.CLcs.LG2026-01被引 9

无需水印或模型信息,精准识别大模型生成文本

Detecting LLM-Generated Text with Performance Guarantees

  • 基于纯文本特征训练分类器,不依赖水印或模型类型
  • 在多个数据集上准确率超现有方法,且控制误报率
  • 支持统计推断,适合需要可信检测的场景

大型语言模型(如GPT、Claude、Gemini和Grok)已深度融入日常应用,广泛用于对话、邮件撰写、教学辅助、编程支持及搜索等任务。然而,其生成高度类人文本的能力引发虚假新闻传播、误导性政府报告及学术不端等风险。为此,我们训练了一个分类器,判断文本是否由大模型生成。该检测器部署于Hugging Face在线CPU平台https://huggingface.co/spaces/stats-powered-ai/StatDetectLLM,具有三项创新:(i)无需水印或特定模型信息;(ii)更有效区分人类与大模型文本;(iii)实现统计推断,弥补当前研究空白。实证结果表明,本方法在分类准确率上优于现有检测器,同时保持类型I误差控制、高统计功效与计算效率。

原文摘要 · Abstract (English)

Large language models (LLMs) such as GPT, Claude, Gemini, and Grok have been deeply integrated into our daily life. They now support a wide range of tasks -- from dialogue and email drafting to assisting with teaching and coding, serving as search engines, and much more. However, their ability to produce highly human-like text raises serious concerns, including the spread of fake news, the generation of misleading governmental reports, and academic misconduct. To address this practical problem, we train a classifier to determine whether a piece of text is authored by an LLM or a human. Our detector is deployed on an online CPU-based platform https://huggingface.co/spaces/stats-powered-ai/StatDetectLLM, and contains three novelties over existing detectors: (i) it does not rely on auxiliary information, such as watermarks or knowledge of the specific LLM used to generate the text; (ii) it more effectively distinguishes between human- and LLM-authored text; and (iii) it enables statistical inference, which is largely absent in the current literature. Empirically, our classifier achieves higher classification accuracy compared to existing detectors, while maintaining type-I error control, high statistical power, and computational efficiency.

文本检测大模型安全统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。