arXiv:2505.21800cs.LGcs.CL2025-05被引 3

发现大模型对真假判断有多个方向的内部表征。

From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs

  • 用多维锥体建模大模型的真假判断机制。
  • 干预锥体可稳定翻转模型对事实的回答。
  • 适用于探究模型内在逻辑,适合研究者使用。

大语言模型(LLMs)具备强大的对话能力,但常生成错误信息。已有研究表明,简单命题的真理性可在模型内部激活中以单一线性方向表示,但这可能无法完整刻画其底层几何结构。本文将近期用于建模拒绝行为的「锥体框架」拓展至真理领域,识别出跨多种大模型家族、因果影响真/假判断的多维锥体。研究通过三条证据支持:(i) 因果干预可稳定翻转模型对事实陈述的响应;(ii) 学习到的锥体在不同模型架构间具有泛化能力;(iii) 锥体干预能保持模型无关行为不变。结果揭示了大模型中简单真/假命题背后更丰富的多向结构,并凸显概念锥体作为探测抽象行为的有力工具。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit strong conversational abilities but often generate falsehoods. Prior work suggests that the truthfulness of simple propositions can be represented as a single linear direction in a model's internal activations, but this may not fully capture its underlying geometry. In this work, we extend the concept cone framework, recently introduced for modeling refusal, to the domain of truth. We identify multi-dimensional cones that causally mediate truth-related behavior across multiple LLM families. Our results are supported by three lines of evidence: (i) causal interventions reliably flip model responses to factual statements, (ii) learned cones generalize across model architectures, and (iii) cone-based interventions preserve unrelated model behavior. These findings reveal the richer, multidirectional structure governing simple true/false propositions in LLMs and highlight concept cones as a promising tool for probing abstract behaviors.

大模型真相判断内部表征锥体模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。