arXiv:2606.06315cs.AI2026-06

让大模型生成内容自带指纹,可精准识别来源且不影响质量

LLM Self-Recognition: Steering and Retrieving Activation Signatures

  • 用随机稀疏向量干预生成过程,植入可检测的激活指纹
  • 在多场景下检测准确率超98%,且不降低文本质量
  • 无需外部水印,利用模型自身结构实现内容溯源,适合内容监管

近期可解释性研究发现,大语言模型(LLMs)在其生成文本中隐式编码了自我识别信号。我们证明该能力在低熵场景下依然可靠,并可通过定向干预增强。通过在生成过程中对内部残差流施加随机稀疏向量,可生成可检测的指纹,实现对特定LLM生成内容的归属判定。该信号可从作为检测器的LLM激活中恢复,在多个检测设置下准确率超过98%,同时保持生成文本质量。随着人工智能生成内容泛滥,该方法提供了一种实用替代方案:利用模型自然表示结构进行内容溯源,而非外部嵌入信号。主要贡献包括:(i) 确立了LLMs可靠的自识别能力;(ii) 提出简单可控的转向机制,实现多模型识别且无质量损失;(iii) 证实激活空间存在可挖掘结构,可用于编码信号而无语义干扰。

原文摘要 · Abstract (English)

Recent advances in interpretability suggest that large language models (LLMs) implicitly encode signals in their generated text that enable self-recognition of their outputs. We demonstrate that this capability is reliable, even in low-entropy scenarios, and that it can be amplified through targeted intervention. By steering the internal residual stream during generation with a random sparse vector, we create a detectable fingerprint that enables attribution of a given text to a specific LLM. This signal is recoverable from the activations of an LLM used as a detector, achieving over 98% accuracy across multiple detection settings while preserving the quality of generated text. As AI-generated content proliferates, this approach offers a practical alternative to traditional detectors by leveraging the model's natural representation structure for attribution rather than embedding a signal externally. Our contributions include: (i) establishing reliable self-recognition capabilities in LLMs, (ii) a simple steering mechanism enabling multi-LLM identification with no quality degradation, (iii) demonstrating that activation spaces contain exploitable structure for encoding signals without semantic interference.

模型溯源自识别可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。