arXiv:2605.12890stat.APcs.LG2026-05

通过引导隐藏状态提升大模型生成文本检测能力

Steer-to-Detect: Probing Hidden Representations for Detection of LLM-Generated Texts

论文配图:Steer-to-Detect: Probing Hidden Representations for Detection of LLM-Generated Texts
图 1 · 摘自论文原文
  • 用引导向量改造模型隐藏状态,增强类别可分性
  • 在多种场景下检测准确率超90%,对抗扰动下仍稳定
  • 首次提供检测过程的理论误差保证,适合安全研究者

大型语言模型(LLMs)的快速发展使机器生成文本与人工写作越来越难以区分。尽管近期研究尝试利用语言模型内部表示挖掘深层检测信号,但原始特征在不同类别间重叠严重,限制了判别能力。为此,我们提出两阶段框架Steer-to-Detect(S2D)。第一阶段,S2D学习一个引导向量注入冻结的观察者模型隐藏状态,生成更具类间分离性的表示;第二阶段,基于引导后的表示进行假设检验完成检测。我们建立了有限样本下对Ⅰ类和Ⅱ类错误的高概率保证,提供了该方法的理论刻画。实验证明,S2D在多种设置下均表现强劲且一致,包括分布外场景和对抗扰动,检测准确率超过90%。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has made machine-generated text increasingly difficult to distinguish from human-written text. While recent studies explore leveraging internal representations of language models to uncover deeper detection signals, these raw features often exhibit substantial overlap between classes, limiting their discriminative power. To address this challenge, we propose Steer-to-Detect (\texttt{S2D}), a two-stage framework for detecting LLM-generated text. In the first stage, \texttt{S2D} learns a steering vector that is injected into the hidden states of a frozen observer LLM, producing representations with improved class separability. In the second stage, detection is performed via a hypothesis testing procedure based on the steered representations. We establish finite-sample, high-probability guarantees for Type I and Type II errors, providing a theoretical characterization of the procedure. Empirically, \texttt{S2D} achieves strong and consistent performance across a range of settings, including out-of-distribution scenarios and adversarial perturbations.

文本检测大模型安全隐空间分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。