arXiv:2604.03754cs.CLcs.AI2026-04被引 2

揭示大模型中真理方向的局限性,发现其依赖层、任务类型和指令设计。

Testing the Limits of Truth Directions in LLMs

论文配图:Testing the Limits of Truth Directions in LLMs
图 1 · 摘自论文原文
  • 真理方向随模型层数变化显著,需多层探测才能理解其本质
  • 事实类任务在早期层出现真理方向,推理类任务则在后期层显现
  • 指令设计会显著影响真理方向的泛化能力,影响模型判断

大语言模型(LLMs)的激活空间中存在线性真理方向,用于编码陈述的真实性。以往研究认为该方向具有普遍性,但近期工作质疑其跨场景泛化能力。本文揭示了真理方向普遍性未被充分认识的若干限制:首先,真理方向具有显著的层依赖性,全面理解需在多层进行探测;其次,其出现位置取决于任务类型——事实类任务在早期层显现,推理类任务在后期层出现,且性能随任务复杂度变化;最后,模型指令对真理方向影响显著,简单正确性评估指令会削弱真理探针的泛化能力。结果表明,真理方向的普遍性远低于先前认知,其表现受模型层、任务难度、任务类型及提示模板的显著影响。

原文摘要 · Abstract (English)

Large language models (LLMs) have been shown to encode truth of statements in their activation space along a linear truth direction. Previous studies have argued that these directions are universal in certain aspects, while more recent work has questioned this conclusion drawing on limited generalization across some settings. In this work, we identify a number of limits of truth-direction universality that have not been previously understood. We first show that truth directions are highly layer-dependent, and that a full understanding of universality requires probing at many layers in the model. We then show that truth directions depend heavily on task type, emerging in earlier layers for factual and later layers for reasoning tasks; they also vary in performance across levels of task complexity. Finally, we show that model instructions dramatically affect truth directions; simple correctness evaluation instructions significantly affect the generalization ability of truth probes. Our findings indicate that universality claims for truth directions are more limited than previously known, with significant differences observable for various model layers, task difficulties, task types, and prompt templates.

大模型真理方向可解释性任务类型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。