arXiv:2507.22149cs.AIcs.LG2025-07EMNLP被引 8

揭示大模型在欺骗指令下内部表征的翻转机制

When Truthful Representations Flip Under Deceptive Instructions?

  • 通过线性探测发现表征可预测输出真假
  • 欺骗指令引发早期到中期层显著表征变化
  • 识别出对欺骗敏感的特异性神经特征

大型语言模型(LLMs)易受恶意构造指令影响,生成欺骗性回应,带来安全风险。现有研究多关注输出层面,缺乏对内部表征如何变化的深入理解。本文以Llama-3.1-8B-Instruct和Gemma-2-9B-Instruct为对象,在事实验证任务中分析其内部表征在欺骗与真实/中性指令下的差异。结果表明,无论指令类型如何,模型输出真/假均可通过线性探测从内部表征中准确预测。进一步利用稀疏自编码器(SAEs)发现,欺骗指令导致显著的表征偏移,而真实与中性指令的表征相似,该变化集中于早期至中期层,并在复杂数据集上仍可检测。我们还定位了对欺骗指令高度敏感的特定SAE特征,并通过可视化确认存在独立的诚实与欺骗表征子空间。研究揭示了欺骗行为在特征与层面上的可辨识信号,为模型检测与控制提供新思路。

原文摘要 · Abstract (English)

Large language models (LLMs) tend to follow maliciously crafted instructions to generate deceptive responses, posing safety challenges. How deceptive instructions alter the internal representations of LLM compared to truthful ones remains poorly understood beyond output analysis. To bridge this gap, we investigate when and how these representations ``flip'', such as from truthful to deceptive, under deceptive versus truthful/neutral instructions. Analyzing the internal representations of Llama-3.1-8B-Instruct and Gemma-2-9B-Instruct on a factual verification task, we find the model's instructed True/False output is predictable via linear probes across all conditions based on the internal representation. Further, we use Sparse Autoencoders (SAEs) to show that the Deceptive instructions induce significant representational shifts compared to Truthful/Neutral representations (which are similar), concentrated in early-to-mid layers and detectable even on complex datasets. We also identify specific SAE features highly sensitive to deceptive instruction and use targeted visualizations to confirm distinct truthful/deceptive representational subspaces. % Our analysis pinpoints layer-wise and feature-level correlates of instructed dishonesty, offering insights for LLM detection and control. Our findings expose feature- and layer-level signatures of deception, offering new insights for detecting and mitigating instructed dishonesty in LLMs.

大模型安全表征分析欺骗检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。