arXiv:2506.00823cs.CL2025-06ACL被引 19

发现大模型的真值方向在逻辑变换中具有泛化能力,可提升回答可信度。

Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks

  • 用简单探针识别真值方向,无需复杂方法
  • 真值方向在逻辑否定等变换中仍保持有效
  • 适合用于筛选可信答案,增强用户信任

大型语言模型(LLMs)在海量数据上训练,蕴含丰富世界知识,但常输出自信的错误内容。早期研究提出‘真值方向’这一线性特征,可可靠判断输出是否真实。本文探讨三个开放问题:(i)LLMs是否普遍具备一致的真值方向;(ii)是否需要高阶探测技术识别真值方向;(iii)真值方向在不同情境下的泛化能力如何。结果表明,并非所有模型都表现出一致真值方向,能力更强的模型在逻辑否定场景中表现更优。此外,基于陈述性事实训练的真值探针,可有效泛化至逻辑变换、问答任务、上下文学习及外部知识源。最后,我们展示了真值探针在选择性问答中的实际应用潜力,有助于提升用户对模型输出的信任。代码已开源于 https://github.com/colored-dye/truthfulness_probe_generalization。

原文摘要 · Abstract (English)

Large language models (LLMs) are trained on extensive datasets that encapsulate substantial world knowledge. However, their outputs often include confidently stated inaccuracies. Earlier works suggest that LLMs encode truthfulness as a distinct linear feature, termed the "truth direction", which can classify truthfulness reliably. We address several open questions about the truth direction: (i) whether LLMs universally exhibit consistent truth directions; (ii) whether sophisticated probing techniques are necessary to identify truth directions; and (iii) how the truth direction generalizes across diverse contexts. Our findings reveal that not all LLMs exhibit consistent truth directions, with stronger representations observed in more capable models, particularly in the context of logical negation. Additionally, we demonstrate that truthfulness probes trained on declarative atomic statements can generalize effectively to logical transformations, question-answering tasks, in-context learning, and external knowledge sources. Finally, we explore the practical application of truthfulness probes in selective question-answering, illustrating their potential to improve user trust in LLM outputs. These results advance our understanding of truth directions and provide new insights into the internal representations of LLM beliefs. Our code is public at https://github.com/colored-dye/truthfulness_probe_generalization

大模型真值方向可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。