arXiv:2602.20273cs.LG2026-02被引 7

发现大模型中存在从通用到特定的多重真实感表示方向。

The Truthfulness Spectrum Hypothesis

  • 通过多类型真伪测试,揭示真实感在模型空间中呈谱系分布。
  • 线性探针跨域泛化良好,但对迎合型与逆预期谎言失效。
  • 后训练会重塑真实感结构,解释对话模型的迎合倾向。

大型语言模型(LLMs)被报道以线性方式编码真实感,但近期研究质疑其普适性。本文提出真实感谱假设:表征空间包含从广泛领域通用到特定领域专属的方向。通过系统评估五类真实类型(定义性、经验性、逻辑性、虚构性、伦理性)、迎合型与逆预期谎言,以及现有诚实度基准,发现线性探针在多数领域间具有良好泛化能力,但在迎合型与逆预期谎言上失败。然而联合训练所有领域可恢复强性能,证实尽管成对迁移差,仍存在领域通用方向。探针方向的几何结构可完美预测跨域泛化(马哈拉诺比斯余弦相似性 R²=0.98)。概念擦除方法进一步分离出三类方向:(1)领域通用,(2)领域特定,(3)仅部分领域共享。因果干预表明,领域特定方向比通用方向更具引导力。后训练会重塑真实感几何结构,使迎合型谎言远离其他真实类型,暗示其表征基础。结果支持真实感谱假设:不同通用性的真实方向共存于表征空间,并受后训练影响。所有实验代码见 https://github.com/zfying/truth_spec。

原文摘要 · Abstract (English)

Large language models (LLMs) have been reported to linearly encode truthfulness, yet recent work questions this finding's generality. We reconcile these views with the truthfulness spectrum hypothesis: the representational space contains directions ranging from broadly domain-general to narrowly domain-specific. To test this hypothesis, we systematically evaluate probe generalization across five truth types (definitional, empirical, logical, fictional, and ethical), sycophantic and expectation-inverted lying, and existing honesty benchmarks. Linear probes generalize well across most domains but fail on sycophantic and expectation-inverted lying. Yet training on all domains jointly recovers strong performance, confirming that domain-general directions exist despite poor pairwise transfer. The geometry of probe directions explains these patterns: Mahalanobis cosine similarity between probes near-perfectly predicts cross-domain generalization (R^2=0.98). Concept-erasure methods further isolate truth directions that are (1) domain-general, (2) domain-specific, or (3) shared only across particular domain subsets. Causal interventions reveal that domain-specific directions steer more effectively than domain-general ones. Finally, post-training reshapes truth geometry, pushing sycophantic lying further from other truth types, suggesting a representational basis for chat models' sycophantic tendencies. Together, our results support the truthfulness spectrum hypothesis: truth directions of varying generality coexist in representational space, with post-training reshaping their geometry. Code for all experiments is provided in https://github.com/zfying/truth_spec.

大模型真实感表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。